Local illumination compensation method, video coding method, apparatus and system
The LIC method addresses illumination changes in video coding by using adjacent sample analysis to determine LIC modes, improving coding efficiency and reducing residual information.
Patent Information
- Application Number
- JP2024572725
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-06-20
- Publication Date
- 2025-07-03
AI Technical Summary
The existing inter prediction methods in video coding technologies, such as H.266/Multipurpose Video Coding (VVC), struggle to effectively handle illumination changes across frames, leading to increased residual information and reduced coding efficiency.
Implement a local illumination compensation (LIC) method that analyzes the availability of adjacent reconstructed samples to determine the LIC mode for inter prediction, using a fitting model to correct illumination differences and reduce residual information.
The LIC method improves coding efficiency by reducing residual information and bit overhead, enhancing the overall prediction effect and coding performance without increasing complexity.
Smart Images

Figure 2025520355000001_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to video technology, and more specifically, but not limited thereto, to a local illumination compensation method, a video coding method, apparatus, and system.
Background Art
[0002] Currently, a block-based hybrid coding framework is used in general-purpose video coding standards such as H.266 / Multipurpose Video Coding (VVC). Each image (frame) in a video is divided into the largest coding units (LCUs) of the same size of a square (e.g., 128×128, 64×64, etc.). Each LCU can also be divided into rectangular coding units (CUs) based on rules. A CU can further be divided into a prediction unit (PU) and a transform unit (TU). The hybrid coding framework can include modules such as prediction, transform, quantization, entropy coding, and in-loop filter. The prediction module can include intra prediction and inter prediction. Inter prediction can include motion estimation and motion compensation. Since there is a strong correlation between adjacent samples in one image within a video, in video coding technology, the intra prediction method is used to eliminate the spatial redundancy between adjacent samples. Since there is a strong similarity between adjacent images in a video, in video coding technology, the inter prediction method can be used to eliminate the temporal redundancy between adjacent images and improve the coding efficiency. However, it is necessary to improve the current inter prediction method to obtain higher coding efficiency. Summary of the Invention
[0003] The following is an overview of the subject matter described in detail in this specification. This overview is not intended to limit the scope of protection of the claims.
[0004] One embodiment of the present disclosure provides a local illumination compensation (LIC) method applied to the decoding side. The LIC method includes the following. Analyze the LIC usage flag of the current block. If it is determined that LIC is used for the current block based on the LIC usage flag, determine the LIC mode used for the current block based on the availability of adjacent reconstructed samples of the current block. When performing inter prediction on the current block, use LIC based on the determined LIC mode.
[0005] One embodiment of the present disclosure further provides a video decoding method. The video decoding method includes the following. If it is determined that a first inter mode is used for the current block, obtain syntax elements at the sequence level and syntax elements at the picture level related to the use of LIC. If it is determined that the use of LIC is permitted when performing inter prediction on the current block based on the syntax elements at the sequence level and the syntax elements at the picture level, use LIC when performing inter prediction on the current block based on the LIC method applied to the decoding side described in any one of the embodiments of the present disclosure.
[0006] One embodiment of the present disclosure further provides a local illumination compensation (LIC) method applied to the encoding side. The LIC method includes the following. If the use of LIC is permitted when performing inter prediction on the current block, determine the LIC mode used when trial-encoding the current block based on the availability of adjacent reconstructed samples of the current block. Trial-encode the current block using the determined LIC mode to calculate the rate-distortion cost, and based on the rate-distortion cost, determine whether to use LIC when performing inter prediction on the current block.
[0007] One embodiment of the present disclosure further provides a video encoding method. The video encoding method includes the following. When a first inter mode is used for the current block, obtain sequence-level syntax elements and picture-level syntax elements related to local illumination compensation (LIC). Based on the sequence-level syntax elements and the picture-level syntax elements, when it is determined that the use of LIC is permitted when performing inter prediction on the current block, use LIC when performing inter prediction on the current block based on the LIC method applied to the encoding side described in any one of the embodiments of the present disclosure.
[0008] One embodiment of the present disclosure further provides a local illumination compensation (LIC) device. The LIC device includes a processor and a memory storing a computer program. When the processor executes the computer program, it enables the execution of the LIC method described in any one of the embodiments of the present disclosure.
[0009] One embodiment of the present disclosure further provides a video decoding device. The video decoding device includes a processor and a memory storing a computer program. When the processor executes the computer program, it enables the execution of the video decoding method described in any one of the embodiments of the present disclosure.
[0010] One embodiment of the present disclosure further provides a video encoding device. The video encoding device includes a processor and a memory storing a computer program. When the processor executes the computer program, it enables the execution of the video encoding method described in any one of the embodiments of the present disclosure.
[0011] One embodiment of the present disclosure further provides a video coding system. The video coding system includes a video encoding device described in any one of the embodiments of the present disclosure and a video decoding device described in any one of the embodiments of the present disclosure.
[0012] One embodiment of the present disclosure further provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium is configured to store a computer program. When the computer program is executed by a processor, the LIC method described in any one of the embodiments of the present disclosure, or the video decoding method described in any one of the embodiments of the present disclosure, or the video encoding method described in any one of the embodiments of the present disclosure is executed.
[0013] One embodiment of the present disclosure further provides a bitstream. The bitstream is generated based on the video encoding method described in any one of the embodiments of the present disclosure.
[0014] After reading and understanding the accompanying drawings and the detailed description, other aspects can be understood.
Brief Description of the Drawings
[0015] The accompanying drawings are provided for the purpose of understanding the embodiments of the present disclosure, form a part of this specification, and are used to explain the technical solutions of the present disclosure together with the embodiments of the present disclosure, and do not constitute a limitation of the technical solutions of the present disclosure.
Figure 1A
Figure 1B
Figure 1C
Figure 2
Figure 3A
Figure 3B
Figure 4
Figure 5A
Figure 5B
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
[0016] This disclosure describes a plurality of embodiments, but the description is illustrative and not limiting, and it is clear to those skilled in the art that more examples and embodiments may be included within the scope of the embodiments described in this disclosure.
[0017] In the description of this disclosure, terms such as "exemplary" or "for example" mean "by way of example, illustration, explanation". In this disclosure, no embodiment described as "for example" or "exemplary" should be construed as being superior to other embodiments. The term "and / or" in this specification is used to describe the relationship of related objects and indicates that there are three types of relationships. For example, in the case of A and / or B, it indicates three situations: only A exists, A and B exist simultaneously, and only B exists. "Plurality" means two or more. Also, in order to clearly describe the technical solutions of the embodiments of this disclosure, terms such as "first", "second", etc. are used to distinguish those with substantially the same or similar functions and roles. Those skilled in the art can understand the following. Terms such as "first", "second", etc. do not limit the number or execution order, and terms such as "first", "second", etc. do not necessarily limit that they are different.
[0018] In the description of representative exemplary embodiments, this specification may present a method and / or process as a specific sequence of steps. However, the method or process does not depend on the specific sequence of steps described in this specification, and the method or process should not be limited to the specific sequence of steps described. As understood by those skilled in the art, other orders of steps are also possible. Therefore, the specific order of steps described in the specification should not be construed as a limitation of the claims. Furthermore, the claims corresponding to the method and / or process should not be limited to being executed in the described order. Those skilled in the art can easily understand that these orders may change and still fall within the spirit and scope of the embodiments of this disclosure.
[0019] The local illuminance compensation (LIC) method and video coding method according to embodiments of the present disclosure can be applied to various video coding standards. Examples include H.264 / Advanced Video Coding (AVC), H.265 / High Efficiency Video Coding (HEVC), H.266 / Multipurpose Video Coding (VVC), Audio Video Standard (AVS), Moving Picture Experts Group (MPEG), standards created by the Alliance for Open Media (AOM), the Joint Video Exploration Team (JVET), and extensions of these standards, or any other customized standards.
[0020] FIG. 1A is a schematic diagram showing a codec system that can be used in an embodiment of the present disclosure. As shown in the figure, the system is divided into an encoding-side device 1 and a decoding-side device 2. The encoding-side device 1 generates a bitstream. The decoding-side device 2 can decode the bitstream. The decoding-side device 2 can receive the bitstream from the encoding-side device 1 via link 3. Link 3 includes one or more media or devices that can move the bitstream from the encoding-side device 1 to the decoding-side device 2. In one example, link 3 includes one or more communication media that enable the encoding-side device 1 to directly transmit the bitstream to the decoding-side device 2. The encoding-side device 1 can modulate the bitstream according to a communication standard (e.g., a wireless communication protocol) and transmit the modulated bitstream to the decoding-side device 2. The one or more communication media include wireless communication media and / or wired communication media and can form part of a packet-based network. In another example, the bitstream can also be output from the output interface 15 to a storage device. The decoding-side device 2 can read the data stored in the storage device via streaming or downloading.
[0021] As shown in FIG. 1A, the encoding-side device 1 includes a data source 11, a video encoding device 13, and an output interface 15. The data source 11 includes a video capture device (e.g., a camera), an archive containing previously captured data, a feeding interface for receiving data from a content provider, a computer graphics system for generating data, or a combination of these sources. The video encoding device 13 can encode the data from the data source 11 and output it to the output interface 15. The output interface 15 can include at least one of a regulator, a modem, and a transmitter. The decoding-side device 2 includes an input interface 21, a video decoding device 23, and a display device 25. The input interface 21 includes at least one of a receiver and a modem. The input interface 21 can receive a bitstream via link 3 or from a storage device. The video decoding device 23 decodes the received bitstream. The display device 25 is used to display the decoded data. The display device 25 may be integrated with other components of the decoding-side device 2 or provided separately. The decoding side may not include the display device 25. In other examples, the decoding side may include other devices or equipment to which the decoded data is applicable.
[0022] Based on the video coding system shown in FIG. 1A, various video coding methods can be used to achieve video compression and decompression.
[0023] FIG. 1B is a block diagram of an exemplary video encoding apparatus that can be used in embodiments of the present disclosure. As shown in the figure, the video encoding apparatus 1000 includes a prediction unit 1100, a splitting unit 1101, a residual generation unit 1102 (shown in the figure by a circled plus following the splitting unit 1101), a transformation processing unit 1104, a quantization unit 1106, an inverse quantization unit 1108, an inverse transformation processing unit 1110, a reconstruction unit 1112 (shown in the figure by a circled plus following the inverse transformation processing unit 1110), a filter unit 1113, a decoded picture buffer 1114, and an entropy encoding unit 1115. The prediction unit 1100 includes an inter prediction unit 1121 and an intra prediction unit 1126. The decoded picture buffer 1114 may be referred to as a buffer for decoded pictures. The video encoder 20 may also include more, fewer, or different functional components compared to this example. In some cases, the transformation processing unit 1104, the inverse transformation processing unit 1110, etc. may not be included.
[0024] The splitting unit 1101, in cooperation with the prediction unit 1100, splits the received video data into slices, coding tree units (CTUs), or other relatively large units. The video data received by the splitting unit 1101 may be a video sequence including video frames such as I-frames, P-frames, and B-frames.
[0025] The prediction unit 1100 can split a CTU into coding units (CUs) and perform intra prediction coding or inter prediction coding on the CUs. When performing intra prediction and inter prediction on a CU, the CU can be split into one or more prediction units (PUs).
[0026] The inter prediction unit 1121 can perform inter prediction on a PU to generate prediction data of the PU. The prediction data includes a predicted block of the PU, motion information of the PU, and various syntax elements. The inter prediction unit 1121 can include a motion estimation unit and a motion compensation unit. The motion estimation unit can be used for motion estimation to generate a motion vector, and the motion compensation unit can be used to obtain or generate a predicted block based on the motion vector.
[0027] The intra prediction unit 1126 can perform intra prediction on a PU to generate prediction data of the PU. The prediction data of the PU can include a predicted block of the PU and various syntax elements.
[0028] The residual generation unit 1102 can subtract a predicted block of a PU obtained by splitting a CU from the original block of the CU to generate a residual block of the CU.
[0029] The transform processing unit 1104 can split a CU into one or more transform units (TUs). The splitting of the prediction unit and the splitting of the transform unit may be different. The residual block related to a TU is a sub-block obtained by splitting the residual block of the CU. By applying one or more transforms to the residual block related to the TU, a coefficient block related to the TU is generated.
[0030] The quantization unit 1106 can quantize the coefficients in the coefficient block based on a selected quantization parameter (QP). The degree of quantization of the coefficient block can be adjusted by adjusting the QP.
[0031] The inverse quantization unit 1108 and the inverse transform processing unit 1110 can respectively apply inverse quantization and inverse transform to the coefficient block to obtain a reconstructed residual block related to the TU.
[0032] The reconstruction unit 1112 can generate a reconstructed image by adding the reconstruction residual block and the prediction block generated by the prediction unit 1100.
[0033] The filter unit 1113 performs in-loop filtering on the reconstructed image and stores the filtered reconstructed image in the decoded image buffer 1114 as a reference image. The intra prediction unit 1126 can extract a reference image of a block adjacent to the PU from the decoded image buffer 1114 and perform intra prediction. The inter prediction unit 1121 can perform inter prediction on the PU of the current image using the reference image of the image before being cached in the decoded image buffer 1114.
[0034] The entropy encoding unit 1115 performs an entropy encoding operation on the received data (e.g., syntax elements, quantized coefficient blocks, motion information, etc.).
[0035] FIG. 1C is a block diagram of an exemplary video decoding apparatus that can be used in an embodiment of the present disclosure. As shown in the figure, the video decoding apparatus 101 includes an entropy decoding unit 150, a prediction unit 152, an inverse quantization unit 154, an inverse transform processing unit 156, a reconstruction unit 158 (shown in the figure with a circled + following the inverse transform processing unit 155), a filter unit 159, and a decoded image buffer 160. In other embodiments, the video decoder 30 may include more, fewer, or different functional components. In some cases, it may not include the inverse transform processing unit 155, etc.
[0036] The entropy decoding unit 150 can perform entropy decoding on the received bitstream to extract information such as syntax elements, quantized coefficient blocks, and motion information of PUs. The prediction unit 152, inverse quantization unit 154, inverse transform processing unit 156, reconstruction unit 158, and filter unit 159 can all execute corresponding operations based on the syntax elements extracted from the bitstream.
[0037] The inverse quantization unit 154 can inverse-quantize the coefficient blocks related to the quantized TUs.
[0038] The inverse transform processing unit 156 can apply one or more inverse transforms to the inverse-quantized coefficient blocks to generate the reconstruction residual blocks of the TUs.
[0039] The prediction unit 152 includes an inter-prediction unit 162 and an intra-prediction unit 164. When intra-prediction coding is used for the PU, the intra-prediction unit 164 can determine the intra-prediction mode of the PU based on the syntax elements parsed from the bitstream, and perform intra-prediction based on the determined intra-prediction mode and the adjacent reconstructed reference information of the PU obtained from the decoded picture buffer 160 to generate the prediction block of the PU. When inter-prediction coding is used for the PU, the inter-prediction unit 162 can determine one or more reference blocks of the PU based on the motion information of the PU and the corresponding syntax elements, and generate the prediction block of the PU based on the reference blocks obtained from the decoded picture buffer 160.
[0040] The reconstruction unit 158 can obtain the reconstructed image based on the reconstruction residual blocks related to the TUs and the prediction blocks of the PUs generated by the prediction unit 152.
[0041] The filter unit 159 can perform in-loop filtering on the reconstructed image. The filtered reconstructed image is stored in the decoded image buffer 160. The decoded image buffer 160 can provide a reference image for use in subsequent motion compensation, intra prediction, inter prediction, etc., and can also output the filtered reconstructed image as decoded video data for display on a display device.
[0042] Based on the above video encoding device and video decoding device, the following basic coding process can be executed. On the encoding side, one image is divided into blocks, and for the current block, an intra prediction or an inter prediction is performed or other algorithms are used to generate a predicted block of the current block. The predicted block is subtracted from the original block of the current block to obtain a residual block. The residual block is transformed and quantized to obtain quantization coefficients, and the quantization coefficients are entropy-coded to output a bitstream. On the decoding side, an intra prediction or an inter prediction is performed on the current block to generate a predicted block of the current block. On the other hand, the quantization coefficients obtained by analyzing the bitstream are inverse quantized and inverse transformed to obtain a residual block, and the predicted block and the residual block are added together to obtain a reconstructed block. The reconstructed blocks form a reconstructed image. The reconstructed image is loop-filtered based on the image or blocks to obtain the decoded image. On the encoding side as well, in order to obtain the decoded image, processing similar to that on the decoding side is required. The decoded image obtained on the encoding side is usually also called the reconstructed image. The decoded image can be used as a reference image for inter prediction of subsequent images. The block division information determined on the encoding side, and mode information or parameter information such as prediction, transformation, quantization, entropy coding, loop filtering, etc. are signaled to the bitstream as needed. The decoding side analyzes the bitstream or analyzes the existing information to determine the same block division information, mode information, and parameter information such as prediction, transformation, quantization, entropy coding, loop filtering, etc. as on the encoding side. Thereby, it is ensured that the decoded image obtained on the encoding side is the same as the decoded image obtained on the decoding side.
[0043] The above is an example of a block-based hybrid coding framework, but the embodiments of the present disclosure are not limited thereto. With the development of technology, one or more modules within the framework and one or more steps within the process can be replaced or optimized.
[0044] Local illumination compensation (LIC) is one of the inter-prediction coding techniques. In this specification, local illumination compensation is also abbreviated as illumination compensation. LIC affects inter-prediction in a video coding hybrid framework and is used on both the encoding side and the decoding side.
[0045] During inter-prediction coding, based on motion vector (MV) information, a reference block for inter-prediction of the current block can be determined. The reference block is usually obtained from different images (e.g., different frames). The illumination intensity changes in the video content of different images. For example, the illumination intensity increases or decreases over time, or changes due to the shielding of dark clouds or a camera flash. The difference between the image of the previous frame and the image of the subsequent frame mainly lies in the strength of the DC component of the image, which is expressed as a change in luminance, and there is little change in the texture information. In motion estimation and motion compensation in inter-prediction technology, such changes cannot be effectively predicted. Due to the influence of a large DC component value, a large amount of residual information is likely to be encoded during encoding. According to the LIC method, these DC redundant information can be well removed, the luminance change can be accurately predicted, and the corresponding compensation can be performed. Thereby, the residual information can be reduced and the coding efficiency can be improved.
[0046] Taking FIGS. 2A and 2B as an example, the texture information of the two images is almost the same, and the difference between them lies in the change of luminance. Since the image in FIG. 2B is illuminated by a camera flash and is very bright, while the image in FIG. 2A is taken under normal natural light, there is a difference between the two. The burden that this difference imposes on video coding is very large. Assuming that for the current block in the image of FIG. 2A, the corresponding reconstructed block in the image of FIG. 2B is used as the reference block, although they have the same texture information, the overall residual is very large. This is because due to the influence of the flash, an overall offset has occurred in the luminance values of the samples in the image of FIG. 2B, and this offset is included in the calculated residual. If this residual is directly transformed, quantized, and signaled (written) to the bitstream, a large overhead will occur.
[0047] In the first embodiment, a LIC method is provided. According to this method, during digital video coding, the illumination difference is corrected by an illumination compensation model to obtain a predicted block after illumination compensation. Thereby, the influence caused by the flash or illumination change is removed, and the overall prediction effect is improved.
[0048] In this embodiment, the relevance between the adjacent reconstructed samples of the current block and the adjacent reconstructed samples of the reference block is used to fit the relevance of the change between the predicted samples and the reference samples within the coding block. If the upper adjacent reconstructed samples and the left adjacent reconstructed samples of the current block in the current image exist, they can be obtained, and the upper adjacent reconstructed samples and the left adjacent reconstructed samples of the reference block in the reference image can also be obtained. Based on the adjacent reconstructed samples of the current block and the adjacent reconstructed samples of the reference block, modeling can be performed to obtain the corresponding fitting model. The fitting model is called the LIC model, and is also called a linear model or simply a model.
[0049] In this specification, "adjacent" means spatially adjacent, and a reconstructed sample is also referred to as a reference sample, a reconstructed reference sample, a template sample, a template reference sample, etc.
[0050] JPEG2025520355000002.jpg35166
Number
[0051] The update of the predicted value is applicable to both the luminance component and the chroma component.
[0052] Based on the adjacent reconstructed samples of the current block and the adjacent reconstructed samples of the reference block, the scaling parameter a and the offset parameter b are calculated. The reconstructed samples shown in FIG. 3A include the upper adjacent reconstructed sample and the left adjacent reconstructed sample of the reference block, and the reconstructed samples in the reference image can be obtained by interpolating the reconstructed pixels. The reconstructed samples shown in FIG. 3B include the upper adjacent reconstructed sample and the left adjacent reconstructed sample of the current block, and the reconstructed samples in the current image may be the reconstructed pixels in the current image. The adjacent reconstructed samples of the reference block and the adjacent reconstructed samples of the current block in the figure are a single row or column, but the present disclosure is not limited thereto.
[0053] JPEG2025520355000005.jpg22170
Number
[0054] In this specification, the current block may be, for example, the current CU or the current PU in the current image. The reference block is, for example, a reconstructed CU or PU for inter prediction for the current block in the reference image, and the position of the reference block is determined based on the current block and the motion vector. The size of the current block is equal to the size of the reference block.
[0055] Before calculating the LIC model parameters of the current block, determine the number and positions of the reconstruction samples for calculating the model parameters based on the width and height of the current block. In one example, if at least one of the width and height of the current block is equal to 4, select 4 reconstruction samples (i.e., 4 upper adjacent reconstruction samples) from the reconstruction samples adjacent to the upper side of the current block and 4 reconstruction samples (i.e., 4 left adjacent reconstruction samples) from the reconstruction samples adjacent to the left side of the current block to calculate the LIC model parameters. For example, if the width of the current block is 16 and the height is 4, select all 4 left adjacent reconstruction samples of the current block, and then, with 3 as the step size, select 4 from the 16 upper adjacent reconstruction samples of the current block to calculate the LIC model parameters. In another example, if neither the width nor the height of the current block is equal to 4, the number of reconstruction samples selected from the upper side of the current block and the number of reconstruction samples selected from the left side of the current block are determined based on the smaller of the width and height of the current block. The number of reconstruction samples to be selected is the logarithm to the base 2 of the smaller of the width and height of the current block, represented by the formula sample_num = log2(min(width, height)). sample_num is the number of reconstruction samples to be selected, and Width and height are the width and height of the current block respectively. For example, if the smaller of the width and height of the current block is 8, 3 reconstruction samples can be selected from the upper side of the current block and 3 reconstruction samples can be selected from the left side of the current block. The number of adjacent reconstruction samples of the reference block for calculating the LIC model parameters is the same as the number of adjacent reconstruction samples of the current block for calculating the LIC model parameters, and their positions correspond to each other.
[0056] JPEG2025520355000008.jpg130167
[0057] JPEG2025520355000009.jpg47170
[0058] In this embodiment, the LIC model parameters are calculated based on the upper adjacent reconstruction samples and left adjacent reconstruction samples of the current block, and the upper adjacent reconstruction samples and left adjacent reconstruction samples of the reference block. This is called the LIC_TL (LIC_top-and-left) mode, where T represents the top side and L represents the left side. The LIC_TL mode may also be referred to as the original illumination compensation mode.
[0059] In the LIC_TL mode, the LIC model parameters are calculated based on the upper adjacent reconstruction samples and left adjacent reconstruction samples of the current block, and the upper adjacent reconstruction samples and left adjacent reconstruction samples of the reference block. When both the upper adjacent reconstruction sample and the left adjacent reconstruction sample of the current block are available, the upper adjacent reconstruction samples and left adjacent reconstruction samples of the current block and the reference block (including the upper adjacent reconstruction samples and left adjacent reconstruction samples of the current block, and the upper adjacent reconstruction samples and left adjacent reconstruction samples of the reference block) are used to calculate the LIC model parameters. However, in the LIC_TL mode, it is not required that both the upper adjacent reconstruction sample and the left adjacent reconstruction sample of the current block be available. When only the upper adjacent reconstruction sample of the current block is available, in the LIC_TL mode, only the upper adjacent reconstruction samples of the current block and the reference block (including the upper adjacent reconstruction sample of the current block and the upper adjacent reconstruction sample of the reference block) are used to calculate the LIC model parameters. When only the left adjacent reconstruction sample of the current block is available, in the LIC_TL mode, only the left adjacent reconstruction samples of the current block and the reference block (including the left adjacent reconstruction sample of the current block and the left adjacent reconstruction sample of the reference block) are used to calculate the LIC model parameters.
[0060] The LIC method in the LIC_TL mode of this embodiment can be used in normal inter prediction (i.e., inter mode), merge prediction mode (i.e., merge mode), and sub-block mode.
[0061] The LIC method in the LIC_TL mode of this embodiment is only used in the single-frame prediction mode and is prohibited from being used in the multi-frame bi-directional reference mode. Also, there is a coupling relationship between the illumination compensation method of this embodiment and other conventional inter-prediction methods. In the current block, the illumination compensation method cannot be used simultaneously with the Bi-directional Optical Flow (BDOF) technology and the SMVD (symmetric motion vector difference) mode.
[0062] In the second embodiment, two different LIC modes, namely the LIC_T (LIC_top) mode and the LIC_L (LIC_left) mode, are added. These two modes are extensions of the LIC_TL mode and are applied to Advanced Motion Vector Prediction (AMVP). Under generalized test conditions, according to the method of this embodiment, the overall coding performance can be improved by 0.07% without increasing the complexity on the decoder side.
[0063] In the LIC_T mode, the LIC model parameters are calculated using only the upper adjacent reconstructed samples of the current block and the upper adjacent reconstructed samples of the reference block. In the LIC_TL mode, when only the upper adjacent reconstructed samples of the current block are available, the LIC model parameters are calculated using only the upper adjacent reconstructed samples of the current block and the upper adjacent reconstructed samples of the reference block. However, different from the LIC_TL mode, in the LIC_T mode, even when both the upper adjacent reconstructed sample and the left adjacent reconstructed sample of the current block are available, the LIC model parameters are calculated using only the upper adjacent reconstructed samples of the current block and the upper adjacent reconstructed samples of the reference block. Please refer to Figure 5A.
[0064] In the LIC_L mode, the LIC model parameters are calculated using only the left adjacent reconstructed samples of the current block and the left adjacent reconstructed samples of the reference block. Also in the LIC_TL mode, when only the left adjacent reconstructed samples of the current block are available, the LIC model parameters are calculated using only the left adjacent reconstructed samples of the current block and the left adjacent reconstructed samples of the reference block. However, different from the LIC_TL mode, in the LIC_L mode, even when both the upper adjacent reconstructed samples and the left adjacent reconstructed samples of the current block are available, the LIC model parameters are calculated using only the left adjacent reconstructed samples of the current block and the left adjacent reconstructed samples of the reference block. Refer to FIG. 5B.
[0065] However, in the second embodiment, when it is determined that the use of LIC is permitted when performing inter prediction on the current block, trial encoding is performed using each of the three LIC modes to calculate the rate-distortion cost, and the optimal LIC mode is selected. The processing of this method is relatively simple. For example, on the decoding side, during the analysis of the bitstream, since the reconstructed samples of the current block located at the boundary are not checked, there may be data redundancy.
[0066] For example, when LIC is used for the current block located at the top boundary of the image, only the left adjacent reconstruction samples are available. When performing trial encoding using the LIC_TL mode, according to the availability of the reconstruction samples, only the left adjacent reconstruction samples of the current block and the left adjacent reconstruction samples of the reference block are used to calculate the LIC model parameters. When performing trial encoding using the LIC_L mode, only the left adjacent reconstruction samples of the current block and the left adjacent reconstruction samples of the reference block are also used to calculate the LIC model parameters. That is, in this case, the LIC_TL mode and the LIC_L mode are the same, and as a result, the calculations are duplicated, and an extra number of bits representing the LIC_L mode are required in the bitstream. As a result, for the CU located at the boundary and using the illumination compensation technology, more bits are undoubtedly transmitted, increasing the overhead. The case where LIC is used for the current block located at the left boundary of the image is similar to the above, and in this case, the LIC_TL mode and the LIC_T mode are the same.
[0067] Also, in the LIC_TL mode, the number of reconstruction samples selected when the current block is at the boundary is the same as the number of reconstruction samples selected when the current block is not at the boundary, and the number of selected reconstruction samples is determined based on the smaller value of the width and height. If LIC is used for the current block located at the top boundary of the image and the size of the current block is 4×32, with a step size of 8, 4 out of the 32 left adjacent reconstruction samples of the current block are selected to calculate the LIC model parameters. With such a selection method, the step size becomes larger, and the calculation of the model parameters may deviate significantly and become inaccurate, leading to deterioration of the coding performance.
[0068] One embodiment of the present disclosure provides a local illumination compensation (LIC) method. The method is applied to the encoding side and, as shown in FIG. 6, includes the following content.
[0069] Step 110: When the use of LIC is permitted for performing inter prediction on the current block, based on the availability of adjacent reconstruction samples of the current block, determine the LIC mode to be used when trial-encoding the current block.
[0070] In this step, the encoding side traverses the prediction modes. When the first inter mode is used for the current block, the LIC-related syntax elements can be obtained to determine whether the use of LIC is permitted when performing inter prediction on the current block. If the use of LIC is permitted, it is necessary to perform trial encoding and determine whether to use LIC based on the rate-distortion cost.
[0071] Step 120: Use the determined LIC mode to trial-encode the current block to calculate the rate-distortion cost, and based on the rate-distortion cost, determine whether to use LIC when performing inter prediction on the current block.
[0072] In this step, after using the determined LIC mode to trial-encode the current block to calculate the rate-distortion cost, the LIC mode with the minimum rate-distortion cost is regarded as the optimal LIC mode. In this case, the encoding side continues to traverse other inter prediction techniques (also referred to as inter prediction modes) such as BDOF and SMVD, calculate the rate-distortion costs corresponding to these modes, and compare the rate-distortion costs corresponding to these modes with the rate-distortion cost of the optimal LIC mode. If the value of the rate-distortion cost of the optimal LIC mode is the minimum, it is determined that LIC is used when performing inter prediction on the current block. Otherwise, LIC is not used when performing inter prediction on the current block. The "optimal LIC mode" is not limited to using a single LIC mode, and may also refer to a combination of the LIC mode and other inter modes that can be used simultaneously with the LIC mode.
[0073] In an embodiment of the present disclosure, when the use of LIC is permitted for inter prediction of the current block, boundary detection for the current block is added to determine the availability of adjacent reconstruction samples of the current block, and based on the availability of the adjacent reconstruction samples of the current block, the LIC mode used when the current block is trial-encoded is determined. When only the upper adjacent reconstruction sample of the current block is available, or when only the left adjacent reconstruction sample of the current block is available, trial encoding is not performed using each of the three LIC modes, namely the LIC_TL mode, the LIC_T mode, and the LIC_L mode, but trial encoding is performed using only one or two of the above three LIC modes. Thereby, data redundancy is removed, the number of bits that need to be transmitted is reduced, and the overhead is lowered.
[0074] In an exemplary embodiment of the present disclosure, determining the LIC mode used when the current block is trial-encoded based on the availability of the adjacent reconstruction samples of the current block includes performing trial encoding using the LIC_TL mode, the LIC_T mode, and the LIC_L mode when both the upper adjacent reconstruction sample and the left adjacent reconstruction sample of the current block are available.
[0075] In an exemplary embodiment of the present disclosure, when only the reconstruction samples adjacent to one side of the current block are available, trial encoding using one LIC mode is skipped, that is, the LIC_T mode or the LIC_L mode is not used during trial encoding.
[0076] In this embodiment, determining the LIC mode used when trial-encoding the current block based on the availability of adjacent reconstruction samples of the current block includes the following. When only the upper adjacent reconstruction samples of the current block are available, trial-encoding is performed using the LIC_TL mode and the LIC_L mode. When only the left adjacent reconstruction samples of the current block are available, trial-encoding is performed using the LIC_TL mode and the LIC_T mode. In the LIC_T mode, the LIC model parameters are calculated using only the upper adjacent reconstruction samples of the current block and the upper adjacent reconstruction samples of the reference block.
[0077] When only the upper adjacent reconstruction samples of the current block are available, when performing trial-encoding using the LIC_L mode, the left adjacent reconstruction samples of the current block can be obtained by filling and the LIC model parameters can be calculated. During filling, the value of the reconstruction sample closest to the left side of the current block among the upper adjacent reconstruction samples of the current block may be assigned to all the left adjacent reconstruction samples of the current block, or one specific value (for example, the median value of the range of sample values) may be assigned to all the left adjacent reconstruction samples of the current block. When only the left adjacent reconstruction samples of the current block are available, when performing trial-encoding using the LIC_T mode, the upper adjacent reconstruction samples of the current block can be obtained by filling and the LIC model parameters can be calculated. The filling method is similar to the above and will not be described in detail here.
[0078] Perform trial encoding using the above method to calculate the rate-distortion cost. When it is determined that LIC is to be used for inter prediction for the current block based on the rate-distortion cost, the above method further includes the following. Set the LIC usage flag (coding unit level flag) of the current block to true and signal it in the bitstream. When the LIC mode with the minimum rate-distortion cost during trial encoding is the LIC_TL mode, set the LIC extension flag of the current block to false and signal it in the bitstream. When the LIC mode with the minimum rate-distortion cost during trial encoding is the LIC_T mode or the LIC_L mode, set the LIC extension flag of the current block to true and signal it in the bitstream.
[0079] In an example of this embodiment, when the LIC mode with the minimum rate-distortion cost is the LIC_T mode or the LIC_L mode, if both the upper adjacent reconstructed sample and the left adjacent reconstructed sample of the current block are available, encode the mode index of the LIC mode with the minimum rate-distortion cost and signal it in the bitstream. When only the upper adjacent reconstructed sample of the current block is available, or when only the left adjacent reconstructed sample of the current block is available, since trial encoding is not performed using the LIC_T mode and the LIC_L mode simultaneously, there is no need to encode the mode index. On the decoding side, based on the availability of the adjacent reconstructed samples of the current block, the optimal LIC mode (i.e., the LIC mode with the minimum rate-distortion cost) can be determined. Specifically, when only the upper adjacent reconstructed sample of the current block is available, determine the optimal LIC mode as the LIC_T mode. When only the left adjacent reconstructed sample of the current block is available, determine the optimal LIC mode as the LIC_T mode.
[0080] In another exemplary embodiment of the present disclosure, if only the reconstruction samples adjacent to one side of the current block are available, the trial encoding in two LIC modes is skipped, that is, the LIC_T mode and the LIC_L mode are not used simultaneously during the trial encoding.
[0081] In this embodiment, determining the LIC mode used when trial-encoding the current block based on the availability of the adjacent reconstruction samples of the current block includes the following. If only the upper adjacent reconstruction sample of the current block is available, or if only the left adjacent reconstruction sample of the current block is available, it is determined to use the LIC_TL mode when trial-encoding the current block.
[0082] When performing trial encoding by the above method to calculate the rate-distortion cost, and it is determined to use LIC when performing inter prediction on the current block based on the rate-distortion cost, the above method further includes the following. Set the LIC usage flag of the current block to true and signal it to the bitstream. If only the upper adjacent reconstruction sample of the current block is available, or if only the left adjacent reconstruction sample of the current block is available, it is not necessary to encode the LIC extension flag and the mode index of the LIC mode of the current block. That is, in this embodiment, if only the reconstruction samples adjacent to one side of the current block are available, only the LIC usage flag of the current block needs to be encoded, reducing the codeword that needs to be transmitted, and accordingly, the decoding process on the decoding side can be simplified.
[0083] If both the upper adjacent reconstruction sample and the left adjacent reconstruction sample of the current block are available, If the LIC mode with the minimum rate-distortion cost during the trial encoding is the LIC_TL mode, set the LIC extension flag of the current block to false and signal it to the bitstream. If the LIC mode with the minimum rate distortion cost is the LIC_T mode or the LIC_L mode, set the LIC extension flag of the current block to true, signal it in the bitstream, encode the mode index of the LIC mode with the minimum rate distortion cost, and signal it in the bitstream.
[0084] In the LIC method of the above embodiment of the present disclosure, when LIC is permitted to be used for the current block, boundary detection for the current block is added. If only the reconstructed samples adjacent to one side of the current block are available, trial encoding using one or more LIC modes is skipped to simplify the calculation process. Also, it is possible to change the encoding of the syntax elements related to LIC of the current block, such as not encoding the mode index of the LIC mode and not encoding the mode index and the LIC extension flag. Thereby, the coding time can be shortened, the overhead can be reduced, and the coding efficiency can be improved.
[0085] One embodiment of the present disclosure further provides a video encoding method. As shown in FIG. 7, the method includes the following content.
[0086] Step 210: When a first inter mode is used for the current block, obtain the syntax elements at the sequence level and the syntax elements at the picture level related to LIC.
[0087] In this step, the encoding side can traverse the prediction mode. If the first inter mode is used for the current block, the sequence-level LIC enable flag and the picture-level LIC usage flag can be obtained. The sequence-level LIC enable flag is used to indicate whether LIC is permitted to be used in the current encoder, and the picture-level LIC usage flag is used to indicate whether LIC is permitted to be used in the current picture.
[0088] Step 220: Based on the sequence-level syntax element and the picture-level syntax element, if it is determined that the use of LIC is permitted when performing inter prediction on the current block, use LIC when performing inter prediction on the current block based on the LIC method applicable to the encoding side described in any embodiment of the present disclosure.
[0089] In one example, when the sequence-level LIC enable flag is true (i.e., indicating that the use of LIC is permitted) and the picture-level LIC usage flag is true (i.e., indicating that LIC is used), it can be determined that the use of LIC is permitted when performing inter prediction on the current block. If at least one of these two flags is false, it is determined that the use of LIC is not permitted when performing inter prediction on the current block.
[0090] In this specification, the picture-level bit indicating that the use of LIC is permitted is called the LIC usage flag, and the CU-level flag indicating that the use of LIC is permitted is also called the LIC usage flag, and the two are easily distinguishable.
[0091] In the video encoding method of this embodiment, using the LIC method applied to the encoding side of any embodiment of the present disclosure, boundary checking is performed on the current block to determine the availability of adjacent reconstructed samples of the current block. When only the reconstructed samples adjacent to one side of the current block are available, the trial encoding by at least one of the LIC_T mode and the LIC_L mode is skipped, the encoding of LIC-related syntax elements (for example, flags, indexes, etc.) is reduced, the coding time is shortened, the bit overhead is saved, and the coding efficiency can be improved.
[0092] In some embodiments, to use LIC, it is necessary to analyze the prediction mode corresponding to the flag into the normal inter prediction mode. The normal inter prediction mode does not include the merge mode. In exemplary embodiments of the present disclosure, the application of the LIC_T mode and the LIC_L mode is extended to the merge mode. In the merge mode, syntax elements are added to indicate the LIC_T mode or the LIC_L mode. In one example, the LIC method of any embodiment of the present disclosure is applicable to the merge mode in the uni-directional reference mode, but not used simultaneously with the combined inter and intra prediction (CIIP) mode, the geometry partition mode (GPM), or the affine_merge mode in the merge mode in the uni-directional reference mode, and can be used simultaneously with MMVD_Merge (Merge mode with Motion Vector Difference), etc. In another example, the LIC method of any embodiment of the present disclosure is applicable to the merge mode in the bi-directional reference mode (in the bi-directional reference mode, LIC is only applicable to the merge mode), but not used simultaneously with the bi-directional optical flow (BDOF) mode, the decoder-side motion vector derivation mode, or the bi-prediction with CU-level weights (BCW) mode in the merge mode in the bi-directional reference mode, and can be used simultaneously with the regular_Merge, etc.
[0093] In an exemplary embodiment of the present disclosure, during the construction of the motion vector prediction (MVP) list in merge mode, the LIC-related flag (e.g., LIC usage flag) of the current block and the mode index of the LIC mode are inherited from an adjacent block or a corresponding predicted block (the adjacent block for inter prediction represents a spatially adjacent block, and the predicted block is obtained from a reference image and represents a temporal block). Alternatively, the LIC usage flag of the current block is inherited from an adjacent block or a corresponding predicted block. However, only the LIC usage flag is inherited, and the mode index of the LIC mode is not inherited. The mode index of the LIC mode is set to the default mode index. By inheriting the LIC usage flag, on the encoding side, no additional rate distortion optimization (RDO) process is required.
[0094] One embodiment of the present disclosure provides a local illumination compensation (LIC) method. The method is applied to the decoding side and includes the following as shown in FIG. 8.
[0095] Step 310: Analyze the LIC usage flag of the current block.
[0096] In this step, when the decoding side determines that the inter mode is used for the current block, it obtains the sequence-level LIC usage permission flag and the picture-level LIC usage flag. When the LIC usage permission flag is true (indicating that the use of LIC is permitted for the current decoder) and the LIC usage flag is true (indicating that LIC is used), the decoding side analyzes the LIC usage flag of the current block to determine whether LIC is used for the current block.
[0097] Step 320: If it is determined that the LIC is used for the current block based on the LIC usage flag, determine the LIC mode used for the current block based on the availability of adjacent reconstruction samples of the current block.
[0098] In this step, on the encoding side, if only the reconstruction samples adjacent to one side of the current block are available, the trial encoding in the LIC_T mode and / or the LIC_L mode is skipped, and it is not necessary to encode the mode index of the LIC mode or the LIC extension flag. Therefore, on the decoding side, according to the availability of the adjacent reconstruction samples of the current block, the syntax elements related to the LIC are analyzed to determine the LIC mode used for the current block.
[0099] Step 330: When performing inter prediction on the current block, use the LIC based on the determined LIC mode.
[0100] In this step, when performing inter prediction on the current block, select the adjacent reconstruction samples of the current block and the adjacent reconstruction samples of the reference block based on the determined LIC mode to calculate the LIC model parameters. Based on the LIC model parameters, linearly transform the prediction block obtained by performing motion compensation on the current block to obtain the prediction block after LIC. The reconstructed block is obtained by summing the prediction block after LIC and the decoded residual.
[0101] In the embodiments of the present disclosure, during decoding, boundary detection is performed on the current block, and according to the availability of the adjacent reconstruction samples of the current block, the syntax elements related to the LIC are analyzed. If only the reconstruction samples adjacent to one side of the current block are available, the LIC mode used for the current block can be determined with relatively few syntax elements and a simpler analysis process. Thereby, the bit overhead can be saved and the coding efficiency can be improved.
[0102] In an exemplary embodiment of the present disclosure, on the encoding side, when only the reconstructed samples adjacent to one side of the current block are available, corresponding to an embodiment where the trial encoding in the LIC_T mode or the LIC_L mode is skipped and the mode index is not encoded, the analysis process and semantics of this embodiment change accordingly. Determining the LIC mode used for the current block based on the availability of the adjacent reconstructed samples of the current block includes the following. Continue to analyze the LIC extension flag of the current block. If the LIC extension flag is true and only the upper adjacent reconstructed samples of the current block are available, it is determined that the LIC_L mode is used for the current block. If the LIC extension flag is true and only the left adjacent reconstructed samples of the current block are available, it is determined that the LIC_T mode is used for the current block.
[0103] On the encoding side, when only the reconstructed samples adjacent to one side of the current block are available and it is determined that either the LIC_T mode or the LIC_L mode is used, since the side where the adjacent reconstructed samples of the current block are available is related to the LIC mode used, the LIC mode used for the current block can be determined without using the mode index.
[0104] In this embodiment, determining the LIC mode used for the current block based on the availability of the adjacent reconstructed samples of the current block further includes the following. If the LIC extension flag is false, it is determined that the LIC_TL mode is used for the current block. If the LIC extension flag is true and both the upper adjacent reconstruction sample and the left adjacent reconstruction sample of the current block are available, continue to analyze the LIC mode index of the current block, and based on the LIC mode index, determine that either the LIC_T mode or the LIC_L mode is used for the current block.
[0105] In another exemplary embodiment of the present disclosure, on the encoding side, if only the reconstruction sample adjacent to one side of the current block is available, skip the trial encoding using the LIC_T mode and the LIC_L mode, and corresponding to an embodiment where the LIC extension flag (used to indicate whether the LIC_T mode or the LIC_L mode is used) and the mode index are not encoded, the analysis process and semantics of this embodiment change accordingly.
[0106] In this embodiment, determining the LIC mode used for the current block based on the availability of the adjacent reconstruction samples of the current block further includes the following. If both the upper adjacent reconstruction sample and the left adjacent reconstruction sample of the current block are available, continue to analyze the LIC extension flag of the current block. If the LIC extension flag is false, determine that the LIC_TL mode is used for the current block. If the LIC extension flag is true, continue to analyze the LIC mode index of the current block, and based on the LIC mode index, determine that either the LIC_T mode or the LIC_L mode is used for the current block.
[0107] In this embodiment, determining the LIC mode used for the current block based on the availability of adjacent reconstruction samples of the current block further includes the following. If it does not hold that both the upper adjacent reconstruction sample and the left adjacent reconstruction sample of the current block are available, it is not necessary to analyze the LIC extension flag, the LIC extension flag is set to false by default, and it is determined that the LIC_TL mode is used for the current block. On the encoding side, when only the reconstruction samples adjacent to one side of the current block are available, the use of the LIC_T mode and the LIC_L mode is not permitted. Therefore, on the decoding side, when it is determined that LIC is used for the current block and only the reconstruction samples adjacent to one side of the current block are available, it is not necessary to analyze other syntax elements, and it can be directly determined that the LIC_TL mode should be used for the current block.
[0108] In an exemplary embodiment of the present disclosure, using LIC based on the determined LIC mode includes the following. When the LIC_T mode is used for the current block, calculate the LIC model parameters using only the upper adjacent reconstruction sample of the current block and the upper adjacent reconstruction sample of the reference block, and perform a linear transformation on the predicted block of the current block based on the LIC model parameters to obtain the predicted block after LIC. When the LIC_L mode is used for the current block, calculate the LIC model parameters using only the left adjacent reconstruction sample of the current block and the left adjacent reconstruction sample of the reference block, and perform a linear transformation on the predicted block of the current block based on the LIC model parameters to obtain the predicted block after LIC. When the LIC_TL mode is used for the current block, calculate the LIC model parameters based on the upper adjacent reconstruction sample and the left adjacent reconstruction sample of the current block, and the upper adjacent reconstruction sample and the left adjacent reconstruction sample of the reference block, and perform a linear transformation on the predicted block of the current block based on the LIC model parameters to obtain the predicted block after LIC.
[0109] One embodiment of the present disclosure further provides a video decoding method. As shown in FIG. 9, the method includes the following content.
[0110] Step 410: When it is determined that the first inter mode is used for the current block, obtain the syntax elements at the sequence level and the syntax elements at the picture level related to the use of the LIC.
[0111] In this step, the analyzed syntax elements at the sequence level and the syntax elements at the picture level can include the LIC usage permission flag at the sequence level and the LIC usage flag at the picture level.
[0112] Step 420: Based on the syntax elements at the sequence level and the syntax elements at the picture level, when it is determined that the use of the LIC is permitted when performing inter prediction on the current block, use the LIC when performing inter prediction on the current block based on the LIC method applied to the decoding side described in any embodiment of the present disclosure.
[0113] In this step, when both the LIC usage permission flag and the LIC usage flag are true, continue to analyze the LIC usage flag of the current block based on the LIC method applied to the decoding side described in any embodiment of the present disclosure. Based on the LIC usage flag, when it is determined that the LIC is used for the current block, determine the LIC mode used for the current block according to the availability of the adjacent reconstructed samples of the current block. Also, when performing inter prediction on the current block, use the LIC based on the determined LIC mode. In this embodiment, during decoding, boundary detection is performed on the current block, and the syntax elements related to the LIC are analyzed according to the availability of the adjacent reconstructed samples of the current block. Thereby, bit overhead can be saved and coding efficiency can be improved.
[0114] In exemplary embodiments of the present disclosure, the application of the LIC_T mode and the LIC_L mode is extended from the AMVP to the merge mode. In the merge mode, syntax elements are added to indicate the LIC_T mode or the LIC_L mode. In one example, the LIC method of any embodiment of the present disclosure is applied to the merge mode in the unidirectional reference mode, but is not used simultaneously with the combined inter-intra prediction (CIIP) mode, the geometric partitioning mode (GPM), or the affine mode in the merge mode in the unidirectional reference mode, and can be used simultaneously with MMVD_Merge or the like. In another example, the LIC method of any embodiment of the present disclosure is applied to the merge mode in the bidirectional reference mode (in the bidirectional reference mode, LIC is applied only to the merge mode), but is not used simultaneously with the bidirectional optical flow (BDOF) mode, the decoder-side motion vector derivation mode, or the bi-prediction by coding unit level weights (BCW) mode in the merge mode in the bidirectional reference mode, and can be used simultaneously with the regular_Merge or the like.
[0115] When constructing the motion vector prediction (MVP) list in the merge mode, the LIC-related flag (e.g., the LIC usage flag) of the current block and the mode index of the LIC mode are inherited from the adjacent block or the corresponding prediction block. Alternatively, the LIC usage flag of the current block is inherited from the adjacent block or the corresponding prediction block. However, only the LIC usage flag is inherited, and the mode index of the LIC mode is not inherited. The mode index of the LIC mode is set as the default mode index. By inheriting the LIC usage flag, the encoding side does not require an additional rate-distortion optimization process.
[0116] In an embodiment of the present disclosure, for the LIC_TL mode, the method for calculating the number of reconstructed samples to be selected is improved, and based on the boundary detection for the current block, the number of reconstructed samples to be selected is determined.
[0117] JPEG2025520355000010.jpg22170
[0118] JPEG2025520355000011.jpg13170
[0119] In the above formula, SampleNum is the base number of the number of reconstructed samples to be selected, cuWidth and cuHeight are the width and height of the current block respectively, cuAbove indicates that the upper adjacent reconstructed sample of the current block is available, and cuLeft indicates that the left adjacent reconstructed sample of the current block is available.
[0120] In the algorithm before improvement, regardless of whether the current block is at the boundary, the number of reconstructed samples to be selected is calculated based on the smaller of the width and height of the current block. In the improved algorithm, boundary detection is added. When the condition that both the upper adjacent reconstructed sample and the left adjacent reconstructed sample of the current block are available is satisfied, the number of reconstructed samples to be selected is calculated based on the smaller of the width and height of the current block. When that condition is not satisfied, it is determined whether the upper adjacent reconstructed sample of the current block is available. When the upper adjacent reconstructed sample of the current block is available, the number of reconstructed samples to be selected is calculated based on the width of the current block. When the upper adjacent reconstructed sample of the current block is not available, the number of reconstructed samples to be selected is calculated based on the height of the current block.
[0121] The improved algorithm of this embodiment can be combined with the LIC method of any embodiment of the present disclosure, and can also be applied to other LIC methods such as the LIC method that uses only the LIC_TL mode.
[0122] When the improved algorithm of this embodiment is combined with the LIC method applied to the encoding side of the embodiments of the present disclosure, the current block is tentatively encoded using the determined LIC mode as follows. When tentatively encoding the current block using the LIC_TL mode, if only the upper adjacent reconstruction samples of the current block are available, the number of reconstruction samples selected for calculating the LIC model parameters is determined based on the width of the current block. If only the left adjacent reconstruction samples of the current block are available, the number of reconstruction samples selected for calculating the LIC model parameters is determined based on the height of the current block. In one example, when determining the number of reconstruction samples selected for calculating the LIC model parameters based on the width of the current block, the number is the logarithm with base 2 of the width. For example, when the width is 8, 3 reconstruction samples may be selected, but the present disclosure is not limited thereto. Determining the number of reconstruction samples selected for calculating the LIC model parameters based on the height of the current block is similar to the above and will not be described in detail.
[0123] When the improved algorithm of this embodiment is combined with the LIC method applied to the decoding side of the embodiments of the present disclosure, calculating the LIC model parameters based on the upper adjacent reconstruction samples and left adjacent reconstruction samples of the current block and the upper adjacent reconstruction samples and left adjacent reconstruction samples of the reference block includes the following. If only the upper adjacent reconstruction samples of the current block are available, the number of reconstruction samples selected for calculating the LIC model parameters is determined based on the width of the current block. If only the left adjacent reconstruction samples of the current block are available, the number of reconstruction samples selected for calculating the LIC model parameters is determined based on the height of the current block.
[0124] In this embodiment, when trial encoding is performed using the LIC_TL mode and the current block is at the boundary of the image, the number of reconstruction samples selected for calculating the LIC model parameters is optimized, avoiding the step size for the selection of the reconstruction samples being too large, improving the accuracy of the calculation of the model parameters, and improving the coding performance.
[0125] One embodiment of the present disclosure provides a video encoding method and a video decoding method. CU-level boundary detection for the illumination compensation mode is added on both the encoding side and the decoding side. When the current block is at the upper boundary, the encoding side skips the trial encoding in the LIC_L mode. When the current block is at the left boundary, the encoding side skips the trial encoding in the LIC_T mode. Accordingly, when the decoding side analyzes the syntax elements of the current block, it determines whether the current block is at the boundary. If the current block is at the boundary, it is not necessary to analyze the last syntax element (mode index). Otherwise, it is necessary to analyze all the LIC-related syntax elements of the current block.
[0126] In this embodiment, the current CU is taken as an example of the current block. The video encoding method of this embodiment includes the following.
[0127] Step 1: The encoder traverses the prediction modes. When an inter mode that can use LIC is applied to the current CU, the LIC usage permission flag and the image-level LIC usage flag are obtained. The LIC usage permission flag is a sequence-level flag indicating that the current encoder is permitted to use LIC.
[0128] When the LIC usage permission flag is true and the image-level LIC usage flag is true, the encoding side tries the LIC prediction method and executes Step 2.
[0129] When the LIC usage permission flag is false or the LIC usage flag is false, the encoding side cannot attempt the LIC prediction method, skips step 2, and directly executes step 3.
[0130] Step 2: The encoder obtains information on the adjacent reconstruction samples of the current CU. If both the upper adjacent reconstruction sample and the left adjacent reconstruction reference sample of the current CU are available, perform the first-round trial encoding, the second-round trial encoding, and the third-round trial encoding. If only the left adjacent reconstruction sample of the current CU is available, only perform the first-round trial encoding and the second-round trial encoding. If only the upper adjacent reconstruction sample of the current CU is available, only perform the first-round trial encoding and the second-round trial encoding.
[0131] In the first-round trial encoding, the LIC_TL mode is used to select the reconstruction samples for calculating the LIC model parameters. The process includes the following. On the encoding side, the upper adjacent reconstruction sample and the left adjacent reconstruction sample of the current CU are selected. The number selected is determined based on the smaller value of the width and height of the current CU. Also, the upper adjacent reconstruction sample and the left adjacent reconstruction sample of the corresponding CU (i.e., the reference CU) in the reference image are selected. When the current CU is non-boundary, both the upper adjacent reconstruction sample and the left adjacent reconstruction sample of the current CU, and the upper adjacent reconstruction sample and the left adjacent reconstruction sample of the reference CU can be selected. When the current CU is on the boundary, only the upper adjacent reconstruction sample or the left adjacent reconstruction sample of the current CU can be selected, and only the upper adjacent reconstruction sample or the left adjacent reconstruction sample of the reference CU can be selected.
[0132] After the reconstructed sample is selected, the obtained reconstructed sample is modeled by the above linear model calculation method, and the scaling coefficient a and the offset parameter b are calculated. Next, a linear transformation is performed on the motion-compensated prediction block using the scaling coefficient a and the offset parameter b to obtain the prediction block after LIC of the current CU. The difference between this prediction block and the original sample corresponding to the current CU is calculated to obtain the residual of the current CU. Rate-distortion cost values are calculated by operations such as transformation and quantization, and are recorded as cost1.
[0133] In the second trial encoding, the LIC_T mode is used to select the reconstructed sample for calculating the LIC model parameters. The process includes the following. On the encoding side, only the upper adjacent reconstructed sample of the current CU is selected, and only the upper adjacent reconstructed sample of the corresponding CU in the reference image is selected. The processing after the reconstructed sample is selected is the same as that in the first round above, and the rate-distortion cost value is calculated and recorded as cost2.
[0134] In the third trial encoding, the LIC_L mode is used to select the reconstructed sample for calculating the LIC model parameters. The process includes the following. On the encoding side, only the left adjacent reconstructed sample of the current CU is selected, and only the left adjacent reconstructed sample of the corresponding CU in the reference image is selected. The processing after the reconstructed sample is selected is the same as that in the first round above, and the rate-distortion cost value is calculated and recorded as cost3.
[0135] cost1, cost2, and cost3 are compared, and the minimum rate-distortion cost value is recorded as costLic, and the information in the current LIC mode (including the mode index of the LIC mode) is saved. The mode index of the LIC_TL mode corresponding to cost1 is 0, the mode index of the LIC_T mode corresponding to cost2 is 1, and the mode index of the LIC_L mode corresponding to cost3 is 2.
[0136] Step 3: The encoding side continues to traverse other inter-prediction modes, calculates the rate-distortion cost value corresponding to each prediction mode, and selects the prediction mode corresponding to the minimum rate-distortion cost value as the optimal prediction mode of the current CU.
[0137] It can be divided into the following several situations.
[0138] In the first situation, if LIC is allowed to be used for the current CU and costLic is the minimum compared with the rate-distortion cost values of other prediction modes, it is determined that LIC is used for the current CU, the CU-level LIC usage flag of the current CU is set to true, and it is signaled in the bitstream.
[0139] If the current CU is non-boundary, if the mode index of the LIC mode corresponding to costLic is greater than 0 (the mode indexes of the LIC_T mode and the LIC_L mode are greater than 0), the LIC extension flag is set to true and signaled in the bitstream, and the mode index is encoded (e.g., by equiprobable encoding) and signaled in the bitstream. If the mode index of the LIC mode corresponding to costLic is equal to 0 (the mode index of the LIC_TL mode is 0), the LIC extension flag is set to false and signaled in the bitstream.
[0140] If the current CU is boundary, if the mode index of the LIC mode corresponding to costLic is greater than 0, the LIC extension flag is set to true and signaled in the bitstream. If the mode index of the LIC mode corresponding to costLic is equal to 0, the LIC extension flag is set to false and signaled in the bitstream. In this case, there is no need to encode the mode index.
[0141] In the second situation, if the LIC is permitted to be used for the current CU and costLic is not the minimum compared with the rate distortion cost values of other prediction modes, the LIC is not used for the current CU, the CU-level LIC usage flag of the current CU is set to false, and it is signaled in the bitstream.
[0142] In the third situation, if the use of the illumination compensation technique for the current CU is not permitted, information such as other optimal prediction modes of the non-illumination compensation technique is signaled in the bitstream. Detailed description is omitted here.
[0143] In this step, if the mode index of the LIC mode corresponding to costLic is greater than 0, costLic may correspond to the LIC_T mode with a mode index of 1 and may also correspond to the LIC_L mode with a mode index of 2. When encoding the mode index and signaling it in the bitstream, 1 can be subtracted from the mode index. In this way, when the mode index is 1, the encoded mode index becomes 0, and when the mode index is 2, the encoded mode index becomes 1. On the decoding side, when the encoded mode index is parsed as 0, it is determined to use the LIC_T mode. When the encoded mode index is parsed as 1, it is determined to use the LIC_L mode.
[0144] Step 4: After the encoder traverses all CUs, it outputs a bitstream after processing such as entropy coding.
[0145] In this specification, the LIC_TL mode may be referred to as the upper and left LIC modes, the LIC_T mode may be referred to as the upper-only LIC mode, and the LIC_T mode may be referred to as the left-only LIC mode.
[0146] The video decoding method of this embodiment includes the following content.
[0147] Step 1: The decoding side analyzes the type flag at the CU level. If the type flag at the CU level is in inter mode, it obtains the LIC usage permission flag. The LIC usage permission flag is a flag at the sequence level, indicating that the LIC is permitted to be used in the current decoder. If the LIC usage permission flag is true, when decoding each image, it is necessary to analyze the LIC usage flag at the image level.
[0148] If both the LIC usage permission flag at the sequence level and the LIC usage flag at the image level are true, continue to analyze the LIC usage flag (flag at the CU level) of the current CU.
[0149] If the LIC usage flag lic_flag of the current CU is true, continue to analyze the LIC extension flag lic_ext.
[0150] If lic_ext is true and both the upper adjacent reconstructed sample and the left adjacent reconstructed sample of the current CU are available, continue to analyze the mode index lic_idx (encoded mode index) of the LIC mode. If lic_idx is 1, it is determined that the LIC_L mode is used for the current CU. If lic_idx is 0, it is determined that the LIC_T mode is used for the current CU.
[0151] If lic_ext is true and only the upper adjacent reconstructed sample or the left adjacent reconstructed sample of the current CU is available, there is no need to analyze the mode index. If only the upper adjacent reconstructed sample of the current CU is available, it is determined that the LIC_L mode is used for the current CU. If only the left adjacent reconstructed sample of the current CU is available, it is determined that the LIC_T mode is used for the current CU.
[0152] If lic_ext is false, it is determined that the LIC_TL mode is used for the current CU.
[0153] After that, Step 2 is executed.
[0154] If the LIC usage flag lic_flag of the current CU is false, Step 3 is executed.
[0155] If one of the sequence-level LIC usage permission flag and the picture-level LIC usage flag is false, in the current decoding process, it is not necessary to analyze the LIC usage flag of the current CU, the CU-level LIC usage flag defaults to false, and Step 3 is executed.
[0156] Step 2: The encoding side determines the LIC mode to be used based on the index obtained from the analysis in the previous step, and based on the determined LIC mode, selects the adjacent reconstruction samples of the current CU. The number of the selected reconstruction samples depends on the width or height of the current CU. Also, select the adjacent reconstruction samples of the corresponding CU in the reference picture. By the linear model calculation method, model the selected reconstruction samples to calculate the scaling coefficient a and the offset parameter b, linearly transform the motion-compensated prediction block, scale it by a times, and compensate it with b to obtain the prediction block after LIC of the current CU.
[0157] Step 3: Continue to analyze information such as the usage flag or index of other inter-prediction modes, and obtain the final prediction block of the current CU based on the analyzed information.
[0158] Step 4: Analyze the bitstream to obtain the residual information, obtain the temporal residual information (also called the decoded residual or the reconstructed residual) by inverse quantization and inverse transformation, and sum the final prediction block and the temporal residual information to obtain the reconstructed sample block of the current block, that is, the reconstructed block.
[0159] Step 5: Perform techniques such as in-loop filtering on all reconstructed sample blocks to obtain a reconstructed image. The reconstructed image may be output as a video or may be used as a reference for subsequent image decoding.
[0160] In this embodiment or other embodiments, a LIC-related flag being 1 indicates that the flag is true, and a LIC-related flag being 0 indicates that the flag is false, but the present disclosure is not limited thereto.
[0161] In this embodiment, at the CU level, the usage restrictions of LIC can be set. If the product of the current CU's width and height is less than 32, the use of LIC is not permitted.
[0162] For the syntax table at the CU level of this embodiment, please refer to the following.
[0163] JPEG2025520355000012.jpg114159
[0164] sps_lic_enable_flag is the LIC usage permission flag at the sequence level, ph_lic_enable_flag is the LIC usage flag at the picture level, lic_flag is the LIC usage flag of the current CU, lic_ext is the LIC extension flag of the current CU, and lic_index is the encoded mode index of the current CU. Using lic_ext + lic_index, the mode index of the LIC mode used for the current CU before encoding can be restored.
[0165] JVET established a research group for coding models beyond H.266 / VVC and named its model, i.e., the platform test software, the Enhanced Compression Model (ECM). In ECM, based on Version 10.0 of the VVC Test Model (VVC TEST MODEL (VTM) 10.0), updated and more efficient compression algorithms were introduced, and many intra-prediction techniques and inter-prediction techniques were integrated. After applying this embodiment to the latest ECM4.0, a general test for random access (RA) is performed. Taking the class D sequence as an example, the test results are as follows.
[0166] JPEG2025520355000013.jpg35159
[0167] As can be seen from the test data of the class D sequence, compared with the above-mentioned second embodiment, in this embodiment, the chroma component is improved. From the encoding side, when at the boundary, the trial encoding is skipped, resulting in effectively shortening the coding time, with an average coding time reduction of 2% - 3%. Also, in this embodiment, redundant information in the bitstream is removed.
[0168] Another embodiment of the present disclosure provides a video encoding method and a video decoding method. Similarly, CU-level boundary detection for the illumination compensation mode is added on both the encoding side and the decoding side. When the current block is at any boundary, the encoding side permits only the original illumination compensation mode, i.e., the LIC_TL mode, and does not permit trial encoding with the LIC_L mode and the LIC_T mode. When analyzing the LIC-related syntax elements of the current block, the decoding side determines whether the current block is at a boundary. When the current block is at any boundary, the decoding side only needs to analyze the first syntax element, i.e., the LIC usage flag. If the LIC usage flag is true, it can be determined to use the LIC_TL mode. If the LIC usage flag is false, it indicates that LIC is not used when performing inter prediction on the current block. In this embodiment, the current CU is taken as an example of the current block.
[0169] The video encoding method of this embodiment includes the following content.
[0170] Step 1: The encoder traverses the prediction mode. When an inter mode that can use LIC is applicable to the current CU, obtain the LIC usage permission flag and the picture-level LIC usage flag. The LIC usage permission flag is a sequence-level flag, indicating that the use of LIC is permitted for the current encoder.
[0171] When the sequence-level LIC usage permission flag is true and the picture-level LIC usage flag is true, the encoding side tries the LIC prediction method and executes Step 2.
[0172] When at least one of the sequence-level LIC usage permission flag and the picture-level LIC usage flag is false, the encoding side cannot try the LIC prediction method, skips Step 2, and directly executes Step 3.
[0173] Step 2: The encoder acquires information on the adjacent reconstruction samples of the current CU. If both the upper adjacent reconstruction sample and the left adjacent reconstruction reference sample of the current CU are available, the first-round trial encoding, the second-round trial encoding, and the third-round trial encoding are performed. If only the left adjacent reconstruction sample of the current CU is available, or if only the upper adjacent reconstruction sample of the current CU is available, only the first-round trial encoding is performed.
[0174] In the first-round trial encoding, reconstruction samples for calculating LIC model parameters are selected using the LIC_TL mode. The process includes the following. On the encoding side, the upper adjacent reconstruction sample and the left adjacent reconstruction sample of the current CU are selected. The number selected is determined based on the smaller value of the width and height of the current CU. Also, the upper adjacent reconstruction sample and the left adjacent reconstruction sample of the corresponding CU (i.e., the reference CU) in the reference image are selected. If the current CU is non-boundary, both the upper adjacent reconstruction sample and the left adjacent reconstruction sample of the current CU, and the upper adjacent reconstruction sample and the left adjacent reconstruction sample of the reference CU can be selected. If the current CU is at the boundary, only the upper adjacent reconstruction sample or the left adjacent reconstruction sample of the current CU is selected, and only the upper adjacent reconstruction sample or the left adjacent reconstruction sample of the reference CU can be selected.
[0175] After the reconstruction samples are selected, the obtained reconstruction samples are modeled using the above linear model calculation method, and the scaling coefficient a and the offset parameter b are calculated. Next, a linear transformation is performed on the motion-compensated prediction block using the scaling coefficient a and the offset parameter b to obtain the prediction block after LIC of the current CU. The difference between this prediction block and the original sample corresponding to the current CU is calculated to obtain the residual of the current CU. Rate distortion cost values are calculated through operations such as transformation and quantization, and are recorded as cost1.
[0176] In the second-pass trial encoding, the LIC_T mode is used to select reconstruction samples for calculating the LIC model parameters. The process includes the following. On the encoding side, only the upper adjacent reconstruction sample of the current CU is selected, and only the upper adjacent reconstruction sample of the corresponding CU in the reference image is selected. The processing after the reconstruction samples are selected is the same as that in the first pass above, and the rate-distortion cost value is calculated and recorded as cost2.
[0177] In the third-pass trial encoding, the LIC_L mode is used to select reconstruction samples for calculating the LIC model parameters. The process includes the following. On the encoding side, only the left adjacent reconstruction sample of the current CU is selected, and only the left adjacent reconstruction sample of the corresponding CU in the reference image is selected. The processing after the reconstruction samples are selected is the same as that in the first pass above, and the rate-distortion cost value is calculated and recorded as cost3.
[0178] cost1, cost2, and cost3 are compared, and the minimum rate-distortion cost value is recorded as costLic, and the information in the current LIC mode (including the mode index of the LIC mode) is saved. The mode index of the LIC_TL mode corresponding to cost1 is 0, the mode index of the LIC_T mode corresponding to cost2 is 1, and the mode index of the LIC_L mode corresponding to cost3 is 2.
[0179] Step 3: The encoding side continues to traverse other inter-prediction modes, calculates the rate-distortion cost value corresponding to each prediction mode, and selects the prediction mode corresponding to the minimum rate-distortion cost value as the optimal prediction mode of the current CU.
[0180] It can be divided into the following several situations.
[0181] In the first situation, if the LIC is permitted to be used for the current CU and the costLic is the minimum compared with the rate-distortion cost values of other prediction modes, it is determined that the LIC is used for the current CU, the CU-level LIC usage flag of the current CU is set to true, and it is signaled in the bitstream.
[0182] If the current CU is non-boundary and the mode index of the LIC mode corresponding to costLic is greater than 0, the LIC extension flag is set to true and signaled in the bitstream, and the mode index is encoded (e.g., by equiprobable coding) and signaled in the bitstream. If the mode index of the LIC mode corresponding to costLic is equal to 0, the LIC extension flag is set to false and signaled in the bitstream.
[0183] If the current CU is boundary, the LIC_TL mode is used by default, and there is no need to signal other syntax elements in the bitstream.
[0184] In the second situation, if the lighting compensation technology is permitted to be used for the current CU and the costLic is not the minimum compared with the rate-distortion cost values of other prediction modes, the LIC is not used for the current CU, the usage flag of the current CU at the CU level is set to false, and it is signaled in the bitstream.
[0185] In the third situation, if the use of LIC is not permitted for the current CU, information such as the optimal prediction mode in which the LIC is not used is signaled in the bitstream. The detailed description is omitted here.
[0186] Step 4: After the encoder traverses all the CUs, it outputs the bitstream after implementing techniques such as entropy coding.
[0187] The video encoding method of this embodiment includes the following content.
[0188] Step 1: The decoding side analyzes the type flag at the CU level. If the type flag at the CU level is in inter mode, it obtains the LIC usage permission flag. The LIC usage permission flag is a flag at the sequence level, indicating that the LIC is permitted to be used in the current decoder. If the LIC usage permission flag is true, when decoding each image, it is necessary to analyze the LIC usage flag at the image level.
[0189] If both the LIC usage permission flag at the sequence level and the LIC usage flag at the image level are true, continue to analyze the LIC usage flag of the current CU.
[0190] If the LIC usage flag lic_flag of the current CU is true and both the upper adjacent reconstructed sample and the left adjacent reconstructed sample of the current CU are available, continue to analyze the LIC extension flag lic_ext of the current CU.
[0191] If lic_ext is true, continue to analyze the mode index lic_idx (encoded mode index) of the LIC mode. If lic_idx is 1, it is determined that the LIC_L mode is used for the current CU. If lic_idx is 0, it is determined that the LIC_T mode is used for the current CU.
[0192] If lic_ext is false, it is determined that the LIC_TL mode is used for the current CU.
[0193] If the lic_flag of the current CU is true and only the upper adjacent reconstructed sample or the left adjacent reconstructed sample of the current CU is available, it is determined that the LIC_TL mode is used for the current CU.
[0194] After that, Step 2 is executed.
[0195] If the current CU's LIC usage flag is false, step 3 is executed.
[0196] If at least one of the sequence-level LIC usage permission flag and the picture-level LIC usage flag is false, in the current decoding process, it is not necessary to analyze the CU-level LIC usage flag, and the CU-level LIC usage flag defaults to false, and step 3 is executed.
[0197] Step 2: The encoder determines the LIC mode to be used based on the index obtained from the analysis in the previous step, and based on the determined LIC mode, selects the adjacent reconstruction samples of the current CU. The number of reconstructed samples to be selected depends on the width or height of the current CU. Also, select the adjacent reconstruction samples of the corresponding CU in the reference picture. By means of the linear model calculation method, model the selected reconstruction samples to calculate the scaling coefficient a and the offset parameter b. Linearly transform the motion-compensated prediction block, scale it by a factor of a, and compensate it with b to obtain the prediction block after LIC of the current CU.
[0198] Step 3: Continue to analyze information such as the usage flags or indexes of other inter-prediction modes, and obtain the final prediction block of the current CU based on the analyzed information.
[0199] Step 4: Analyze the bitstream to obtain the residual information, and obtain the temporal residual information by inverse quantization and inverse transformation. Sum the final prediction block and the temporal residual information to obtain the reconstructed sample block of the current CU.
[0200] Step 5: Execute techniques such as in-loop filtering on all the reconstructed sample blocks to obtain the reconstructed image. The reconstructed image may be output as a video or may be used as a reference for subsequent decoding.
[0201] In this embodiment, at the CU level, usage restrictions of the LIC can be set. When the product of the current width and height of the CU is less than 32, the use of the LIC is not permitted.
[0202] For the syntax table at the CU level of this embodiment, please refer to the following.
[0203] JPEG2025520355000014.jpg118169
[0204] sps_lic_enable_flag is the LIC usage permission flag at the sequence level, ph_lic_enable_flag is the LIC usage flag at the picture level, lic_flag is the LIC usage flag of the current CU, lic_ext is the LIC extension flag of the current CU, and lic_index is the encoded mode index of the current CU. With lic_ext + lic_index, the mode index of the LIC mode used for the current CU before encoding can be restored.
[0205] One embodiment of the present disclosure further provides a local illumination compensation (LIC) device. As shown in FIG. 10, the LIC device includes a processor 71 and a memory 73 in which a computer program is stored. When the processor 71 executes the computer program, it executes the LIC method described in any one of the embodiments of the present disclosure.
[0206] One embodiment of the present disclosure further provides a video decoding device. As shown in FIG. 10, the video decoding device includes a processor and a memory in which a computer program is stored. When the processor executes the computer program, it executes the video decoding method described in any one of the embodiments of the present disclosure.
[0207] One embodiment of the present disclosure further provides a video encoding device. As shown in FIG. 10, the video encoding device includes a processor and a memory in which a computer program is stored. When the processor executes the computer program, it executes the video encoding method according to any one of the embodiments of the present disclosure.
[0208] The processor in the above embodiment of the present disclosure can be a general-purpose processor including a central processing unit (CPU), a network processor (NP), a microprocessor, etc., or can be other ordinary processors, etc. The processor can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a discrete logic circuit, or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, or a combination of the above components. That is, the processor in the above embodiment can be any processing component or combination of components in the various methods, steps and logic block diagrams disclosed in the embodiments of the present invention. When the embodiment of the present disclosure is implemented partially in software, the instructions used in the software can be stored in a suitable non-volatile computer-readable storage medium, and one or more processors can implement the method of the embodiment of the present disclosure by executing the instructions in hardware.
[0209] One embodiment of the present disclosure further provides a video coding system. The video coding system includes the video encoding device described in any one embodiment of the present disclosure and the video decoding device described in any one embodiment of the present disclosure.
[0210] One embodiment of the present disclosure further provides a non-transitory computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the LIC method described in any one embodiment of the present disclosure is executed, or the video decoding method described in any one embodiment of the present disclosure is executed, or the video encoding method described in any one embodiment of the present disclosure is executed.
[0211] One embodiment of the present disclosure further provides a bitstream. The bitstream is generated based on the video encoding method described in any one embodiment of the present disclosure.
[0212] In one or more of the above exemplary embodiments, the functions described can be implemented by hardware, software, firmware, or any combination thereof. When implemented by software, the functions can be stored on a computer-readable medium as one or more instructions or codes, or can be transmitted via a computer-readable medium and executed by a hardware-based processing unit. The computer-readable medium includes a computer-readable medium that is a tangible medium such as a data storage medium, or can include any communication medium that facilitates transmission of a computer program from one place to another, for example, in accordance with a communication protocol. Thus, the computer-readable medium can typically be either a non-transitory tangible computer-readable storage medium or a communication medium such as a signal or carrier. The data storage medium can be any available medium that can be accessed by one or more computers or one or more processors to search for instructions, codes, and / or data structures for implementing the techniques described in this disclosure. A computer program product can include a computer-readable medium.
[0213] By way of non-limiting example, such a computer-readable storage medium can include random access memory (RAM), read only memory (ROM), electrically erasable programmable ROM (EEPROM), compact disk ROM (CD-ROM) or other optical disk storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory, or any other medium that can store the desired program code in the form of instructions or data structures and is accessible by a computer. Also, any connection can be referred to as a computer-readable storage medium. For example, when transmitting instructions from a website, server or other remote source using coaxial cable, fiber optic cable, twisted-pair cabling, digital subscriber line (DSL), or wireless technologies such as infrared, radio, microwave, etc., coaxial cable, fiber optic cable, twisted-pair cabling, DSL, or wireless technologies such as infrared, radio, microwave, etc. are included in the definition of the medium. However, computer-readable storage media and data storage media do not include connections, carriers, signals, or other transient media, but are non-transitory tangible storage media. As used herein, magnetic disks and optical disks include compact disks (CDs), laser disks, optical disks, digital versatile disks (DVDs), floppy disks, or Blu-ray disks, etc. Magnetic disks typically reproduce data magnetically, and optical disks reproduce data optically using a laser. The above combinations should also be included within the scope of computer-readable media.
[0214] The commands can be executed by one or more processors, such as one or more DSPs, general-purpose microprocessors, ASICs, FPGAs, or other equivalent integrated circuits or discrete logic circuits. Thus, the term "processor" as used herein can refer to any one of the above structures, or any other structure suitable for implementing the techniques described herein. Further, in some embodiments, the functions described herein can be provided within dedicated hardware and / or software modules configured to be used in encoding and decoding, and can also be incorporated into an integrated encoder-decoder. Also, the techniques described herein can be fully realized in one or more circuits or logic elements.
[0215] The technical solutions of the embodiments of the present disclosure can be implemented in various devices or apparatuses including wireless mobile phones, integrated circuits (ICs), or a set of ICs (e.g., a chipset). In the embodiments of the present disclosure, various components, modules, or units are used to emphasize the functions of a device configured to execute the described techniques. It is not necessarily realized by different hardware units. As described above, the various units may be combined in the codec hardware units, or provided in combination with a suitable software and / or firmware in an aggregate of interoperable hardware units (including one or more of the processors described above).
Claims
1. A local illumination compensation (LIC) method applied to the decoding side, comprising: analyzing the LIC usage flag of the current block; when it is determined that LIC is used for the current block based on the LIC usage flag, determining the LIC mode used for the current block based on the availability of adjacent reconstruction samples of the current block; when performing inter prediction on the current block, using LIC based on the determined LIC mode. A local illumination compensation method, characterized by the above.
2. Determining the LIC mode used for the current block based on the availability of adjacent reconstruction samples of the current block includes: continuing to analyze the LIC extension flag of the current block; when the LIC extension flag is true and only the upper adjacent reconstruction sample of the current block is available, determining that the LIC_L (LIC_left) mode is used for the current block, wherein in the LIC_L mode, the LIC model parameters are calculated using only the left adjacent reconstruction sample of the current block and the left adjacent reconstruction sample of the reference block; when the LIC extension flag is true and only the left adjacent reconstruction sample of the current block is available, determining that the LIC_T (LIC_top) mode is used for the current block, wherein in the LIC_T mode, the LIC model parameters are calculated using only the upper adjacent reconstruction sample of the current block and the upper adjacent reconstruction sample of the reference block. A method according to Claim 1, characterized by the above.
3. Determining the LIC mode used for the current block based on the availability of adjacent reconstruction samples of the current block includes: when the LIC extension flag is false, determining that the LIC_TL (LIC_top-and-left) mode is used for the current block, wherein in the LIC_TL mode, the LIC model parameters are calculated based on the upper adjacent reconstruction sample and the left adjacent reconstruction sample of the current block, and the upper adjacent reconstruction sample and the left adjacent reconstruction sample of the reference block. When the LIC extension flag is true and both the upper adjacent reconstructed sample and the left adjacent reconstructed sample of the current block are available, continue to analyze the LIC mode index of the current block, and based on the LIC mode index, determine that either the LIC_T mode or the LIC_L mode is used for the current block, further comprising, The method according to claim 2, characterized in that.
4. Determining the LIC mode used for the current block based on the availability of the adjacent reconstructed samples of the current block is When both the upper adjacent reconstructed sample and the left adjacent reconstructed sample of the current block are available, continue to analyze the LIC extension flag of the current block, When the LIC extension flag is false, determine that the LIC_TL mode is used for the current block, When the LIC extension flag is true, continue to analyze the LIC mode index of the current block, and based on the LIC mode index, determine that either the LIC_T mode or the LIC_L mode is used for the current block, comprising, The method according to claim 1, characterized in that.
5. Determining the LIC mode used for the current block based on the availability of the adjacent reconstructed samples of the current block is When it is not established that both the upper adjacent reconstructed sample and the left adjacent reconstructed sample of the current block are available, it is not necessary to analyze the LIC extension flag, set the LIC extension flag to false by default, and determine that the LIC_TL mode is used for the current block, further comprising, The method according to claim 4, characterized in that.
6. Using LIC based on the determined LIC mode is When the LIC_T mode is used for the current block, calculate the LIC model parameters using only the upper adjacent reconstructed sample of the current block and the upper adjacent reconstructed sample of the reference block, and perform a linear transformation on the predicted block of the current block based on the LIC model parameters to obtain the predicted block after LIC, When the LIC_L mode is used for the current block, calculate the LIC model parameters using only the left adjacent reconstruction samples of the current block and the left adjacent reconstruction samples of the reference block, and perform a linear transformation on the predicted block of the current block based on the LIC model parameters to obtain the predicted block after LIC, and When the LIC_TL mode is used for the current block, calculate the LIC model parameters based on the upper adjacent reconstruction samples and the left adjacent reconstruction samples of the current block and the upper adjacent reconstruction samples and the left adjacent reconstruction samples of the reference block, and perform a linear transformation on the predicted block of the current block based on the LIC model parameters to obtain the predicted block after LIC, and including The method according to claim 1, characterized in that.
7. Calculating the LIC model parameters based on the upper adjacent reconstruction samples and the left adjacent reconstruction samples of the current block and the upper adjacent reconstruction samples and the left adjacent reconstruction samples of the reference block includes When only the upper adjacent reconstruction samples of the current block are available, determining the number of reconstruction samples selected for calculating the LIC model parameters based on the width of the current block, and When only the left adjacent reconstruction samples of the current block are available, determining the number of reconstruction samples selected for calculating the LIC model parameters based on the height of the current block, and including The method according to claim 6, characterized in that.
8. A video decoding method, comprising: When it is determined that the first inter mode is used for the current block, obtaining the syntax elements at the sequence level and the syntax elements at the picture level related to the use of local illumination compensation (LIC); and Based on the syntax elements at the sequence level and the syntax elements at the picture level, when it is determined that the use of LIC is permitted when performing inter prediction on the current block, using LIC when performing inter prediction on the current block based on the method according to any one of claims 1 to 7, and including The video decoding method is characterized in that.
9. The first inter mode is A merge mode in unidirectional reference mode, which is not used simultaneously with a combination of inter prediction and intra prediction (CIIP) mode, a geometric partitioning mode, or an affine mode in the merge mode in unidirectional reference mode, and A merge mode in bidirectional reference mode, which is not used simultaneously with a bidirectional optical flow (BDOF) mode, a decoder-side motion vector derivation mode, or a bi-prediction (BCW) mode based on coding unit level weights in the merge mode in bidirectional reference mode, and at least one of The method according to claim 8, characterized in that.
10. The method During the construction of the motion vector prediction list in merge mode, the LIC usage flag and the mode index of the LIC mode of the current block are inherited from an adjacent block or a corresponding prediction block, or the LIC usage flag of the current block is inherited from an adjacent block or a corresponding prediction block, and the mode index of the LIC_TL (LIC_top-and-left) mode is set as the default mode index, further including The method according to claim 9, characterized in that.
11. A local illumination compensation (LIC) method applied to the encoding side, When the use of LIC is permitted when performing inter prediction on the current block, determining the LIC mode used when trial-encoding the current block based on the availability of adjacent reconstructed samples of the current block; and Trial-encoding the current block using the determined LIC mode to calculate a rate-distortion cost, and determining whether to use LIC when performing inter prediction on the current block based on the rate-distortion cost; and including The local illumination compensation method is characterized in that.
12. Determining the LIC mode used when trial-encoding the current block based on the availability of adjacent reconstructed samples of the current block is When both the upper adjacent reconstruction sample and the left adjacent reconstruction sample of the current block are available, performing trial encoding using the LIC_TL (LIC_top-and-left) mode, the LIC_T (LIC_top) mode, and the LIC_L (LIC_left) mode, including The method according to claim 11, characterized in that
13. Based on the availability of the adjacent reconstruction samples of the current block, determining the LIC mode used when trial-encoding the current block is When only the upper adjacent reconstruction sample of the current block is available, performing trial encoding using the LIC_TL mode and the LIC_L mode. In the LIC_TL mode, the LIC model parameters are calculated based on the upper adjacent reconstruction sample and the left adjacent reconstruction sample of the current block, and the upper adjacent reconstruction sample and the left adjacent reconstruction sample of the reference block. In the LIC_L mode, the LIC model parameters are calculated using only the left adjacent reconstruction sample of the current block and the left adjacent reconstruction sample of the reference block, performing When only the left adjacent reconstruction sample of the current block is available, performing trial encoding using the LIC_TL mode and the LIC_T mode. In the LIC_T mode, the LIC model parameters are calculated using only the upper adjacent reconstruction sample of the current block and the upper adjacent reconstruction sample of the reference block, performing including The method according to claim 11 or 12, characterized in that
14. After determining whether to use LIC when performing inter prediction on the current block based on the rate-distortion cost, the method If it is determined to use LIC, setting the LIC usage flag of the current block to true and signaling it in the bitstream If the LIC mode with the minimum rate-distortion cost during trial encoding is the LIC_TL mode, setting the LIC extension flag of the current block to false and signaling it in the bitstream If the LIC mode with the minimum rate-distortion cost during trial encoding is the LIC_T mode or the LIC_L mode, setting the LIC extension flag of the current block to true and signaling it in the bitstream further including The method according to claim 13, characterized in that.
15. When the LIC mode with the minimum rate distortion cost during trial encoding is the LIC_T mode or the LIC_L mode, the method comprises: When only the upper adjacent reconstruction sample of the current block is available, or when only the left adjacent reconstruction sample of the current block is available, it is not necessary to encode the mode index; When both the upper adjacent reconstruction sample and the left adjacent reconstruction sample of the current block are available, encoding the mode index of the LIC mode with the minimum rate distortion cost and signaling it in the bitstream; further comprising The method according to claim 14, characterized in that.
16. Based on the availability of the adjacent reconstruction samples of the current block, determining the LIC mode used when trial-encoding the current block When only the upper adjacent reconstruction sample of the current block is available, or when only the left adjacent reconstruction sample of the current block is available, determining to use the LIC_TL mode when trial-encoding the current block, including The method according to claim 11 or 12, characterized in that.
17. After determining whether to use LIC when performing inter prediction on the current block based on the rate distortion cost, the method If it is determined to use LIC, setting the LIC usage flag of the current block to true and signaling it in the bitstream; When only the upper adjacent reconstruction sample of the current block is available, or when only the left adjacent reconstruction sample of the current block is available, it is not necessary to encode the LIC extension flag and the mode index of the LIC mode of the current block; further comprising The method according to claim 16, characterized in that.
18. If it is determined to use LIC, the method When both the upper adjacent reconstruction sample and the left adjacent reconstruction sample of the current block are available If the LIC mode with the minimum rate distortion cost during trial encoding is the LIC_TL mode, setting the LIC extension flag of the current block to false and signaling it in the bitstream; If the LIC mode with the minimum rate distortion cost is the LIC_T mode or the LIC_L mode, set the LIC extension flag of the current block to true, signal it in the bitstream, encode the mode index of the LIC mode with the minimum rate distortion cost, and signal it in the bitstream. further comprising The method according to claim 17, characterized in that.
19. When trial-encoding the current block using the determined LIC mode, When trial-encoding the current block using the LIC_TL mode, If only the upper adjacent reconstructed samples of the current block are available, determine the number of reconstructed samples selected for calculating the LIC model parameters based on the width of the current block. If only the left adjacent reconstructed samples of the current block are available, determine the number of reconstructed samples selected for calculating the LIC model parameters based on the height of the current block. including The method according to claim 11, characterized in that.
20. A video encoding method, When a first inter mode is used for the current block, obtain the syntax elements at the sequence level and the syntax elements at the picture level related to local illumination compensation (LIC). Based on the syntax elements at the sequence level and the syntax elements at the picture level, if it is determined that the use of LIC is permitted when performing inter prediction on the current block, use LIC when performing inter prediction on the current block based on the method according to any one of claims 11 to 19. including A video encoding method characterized by that.
21. The first inter mode is The merge mode in the unidirectional reference mode, which is not used simultaneously with the combination of inter prediction and intra prediction (CIIP) mode, geometric partitioning mode, or affine mode in the merge mode in the unidirectional reference mode. A merge mode in a bidirectional reference mode, which is not used simultaneously with a bidirectional optical flow (BDOF) mode, a decoder-side motion vector derivation mode, or a bi-prediction (BCW) mode based on coding unit level weights in the merge mode in the bidirectional reference mode, and including at least one of The method according to claim 20, characterized in that.
22. The method is During the construction of the motion vector prediction list in the merge mode, the LIC usage flag and the mode index of the LIC mode of the current block are inherited from an adjacent block or a corresponding prediction block, or the LIC usage flag of the current block is inherited from an adjacent block or a corresponding prediction block, and the mode index of the LIC_TL (LIC_top-and-left) mode is set as the default mode index, further including The method according to claim 21, characterized in that.
23. A local illumination compensation (LIC) device, comprising a processor and a memory storing a computer program, and when the processor executes the computer program, enabling execution of the local illumination compensation method according to any one of claims 1 to 7, claims 11 to 19, A local illumination compensation device, characterized in that.
24. A video decoding device, comprising a processor and a memory storing a computer program, and when the processor executes the computer program, enabling execution of the video decoding method according to any one of claims 8 to 10, A video decoding device, characterized in that.
25. A video encoding device, comprising a processor and a memory storing a computer program, and when the processor executes the computer program, enabling execution of the video encoding method according to any one of claims 20 to 22, A video encoding device, characterized in that.
26. A video coding system, comprising the video encoding device according to claim 25 and the video decoding device according to claim 24, A video coding system, characterized in that.
27. A non-transitory computer-readable storage medium, The non-transitory computer-readable storage medium is configured to store a computer program, and when the computer program is executed by a processor, the local illumination compensation method according to any one of claims 1 to 7, claims 11 to 19, or the video decoding method according to any one of claims 8 to 10, or the video encoding method according to any one of claims 20 to 22 is executed. A non-transitory computer-readable storage medium, characterized in that.
28. A bitstream, wherein the bitstream is generated based on the video encoding method according to any one of claims 20 to 22. A bitstream, characterized in that.
Citation Information
Patent Citations
Moving image decoding device and moving image encoding device
JP2020195014A
Motion compensation in video encoding and decoding
JP2021523604A
Method, apparatus and computer program for video encoding or decoding
JP2022508161A
Neighboring Sample Selection for Intra Prediction
JP2022521698A
Simplification for cross-component linear model prediction mode
US20210344966A1