Improved mode processing in video encoding and decoding
The refinement step in video encoding and decoding addresses distortion issues by using refinement tables to adjust video signals, enhancing compression efficiency and quality.
Patent Information
- Application Number
- JP2024119599
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-07-25
- Publication Date
- 2025-12-24
- Estimated Expiration
- 2039-06-11
AI Technical Summary
Existing video encoding and decoding technologies face challenges in achieving high compression efficiency while minimizing distortion in reconstructed video signals, particularly when applying improvements that benefit the entire slice or region may degrade local portions.
Implementing a refinement step in the encoding and decoding process that uses refinement tables based on refinement parameters to adjust video signals, specifically through inter-component and intra-component improvements, allowing for reduced distortion and improved compression efficiency.
The refinement process reduces distortion and enhances coding efficiency by applying block-based refinement parameters, resulting in improved video quality for the same bit rate or reduced bit rate for the same quality.
Smart Images

Figure 0007791942000011 
Figure 0007791942000012 
Figure 0007791942000013
Abstract
Description
[Technical Field]
[0001] Technical Field This disclosure includes video encoding and decoding. [Background technology]
[0002] background To achieve high compression efficiency, image and video coding schemes typically use prediction and transformation to exploit spatial and temporal redundancy in video content. Generally, intra- or inter-prediction is used to exploit correlation within or between frames, and then the difference between the original and predicted picture blocks, often denoted as prediction error or prediction residual, is transformed, quantized, and entropy coded. To reconstruct (reconstruct) the video, the compressed data is decoded by the inverse process corresponding to prediction, transformation, quantization, and entropy coding. Summary of the Invention
[0003] overview In general, an example of at least one embodiment of the method includes obtaining (i) a cost for encoding a picture portion based on applying a refinement mode to a block-by-block reconstructed signal, the refinement mode being based on refinement parameters, and (ii) a cost for encoding the picture portion using a mode other than the refinement mode, and encoding the picture portion based on the cost.
[0004] In general, an example of at least one embodiment of an apparatus includes one or more processors configured to obtain a cost for encoding a picture portion based on applying an improvement mode to a block-by-block reconstructed signal, where the improvement mode is based on an improvement parameter, and a cost for encoding the picture portion without using the improvement mode, and to encode the picture portion based on the cost.
[0005] In general, an example of at least one embodiment of a decoding method includes obtaining an indication of an improvement mode that has been applied to a block-by-block reconstructed signal during encoding of the picture portion, the improvement mode being based on an improvement parameter, and decoding the encoded picture portion based on the indication.
[0006] In general, an example of at least one embodiment of an apparatus includes one or more processors configured to obtain an indication of an improvement mode that has been applied to a block-by-block reconstructed signal during encoding of the encoded picture portion, the improvement mode being based on an improvement parameter, and decoding the encoded picture portion based on the indication.
[0007] In general, one example of at least one embodiment includes providing a signal formatted to include data representing an encoded picture portion and data providing an indication of an enhancement mode that has been applied to a block-by-block reconstructed signal based on enhancement parameters during encoding of the encoded picture portion.
[0008] In general, one example of at least one embodiment includes providing a bitstream formatted to include data representing an encoded picture portion and data providing an indication of an enhancement mode that has been applied to a block-by-block reconstructed signal based on enhancement parameters during encoding of the encoded picture portion.
[0009] In general, an example of at least one embodiment includes a computer program product or non-transitory computer-readable storage medium storing computer-readable instructions for causing one or more processors to perform any of the methods described herein.
[0010] Generally, one or more embodiments provide a computer-readable storage medium, e.g., a non-volatile computer-readable storage medium, storing instructions for encoding or decoding video data in accordance with a method or apparatus described herein and / or storing a bitstream generated in accordance with a method or apparatus described herein. One or more of the embodiments herein may also provide methods and apparatus for transmitting or receiving a bitstream generated in accordance with a method or apparatus described herein.
[0011] In general, another example of at least one embodiment includes a device including one or more processors configured to encode or decode picture portions in accordance with a method or apparatus described herein, and including at least one of: (i) an antenna configured to receive a signal, the signal including data representing an image; (ii) a band limiter configured to limit the received signal to a frequency band including the data representing the image; or (iii) a display configured to display the image.
[0012] BRIEF DESCRIPTION OF THE DRAWINGS The present disclosure can be better understood from the following detailed description considered in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 2 is a block diagram illustrating an example embodiment of a video encoder. [Figure 2] FIG. 2 is a block diagram illustrating an example embodiment of a video decoder. [Figure 3] FIG. 2 is a block diagram illustrating another example embodiment of a video encoder. [Figure 4] FIG. 2 is a block diagram illustrating another example embodiment of a video decoder. [Figure 5] 1 is a table illustrating an example embodiment of a signal syntax. [Figure 6] 10 is a graph illustrating features of an embodiment of an improved table. [Figure 7]10 is a table illustrating another example embodiment of a signal syntax. [Figure 8] 10 is a table illustrating an example of features of another embodiment of a signal syntax. [Figure 9] 1 is a flow diagram illustrating an example embodiment of a process for deriving refinement information. [Figure 10] 10 is a table illustrating another example embodiment of a signal syntax. [Figure 11] 10 is a flow diagram illustrating another example embodiment of a process for deriving refinement information. [Figure 12] 1A-1C are two flow charts illustrating an example embodiment of a process for encoding and decoding refinement parameters. [Figure 13] 10 is a table illustrating another example embodiment of a signal syntax. [Figure 14] 1 is a flow diagram illustrating an example embodiment of a process for deriving information for use in one or more embodiments. [Figure 15] 1 is a block diagram illustrating an example of an embodiment of a system that includes video encoding and / or decoding. [Figure 16] 1 is a block diagram illustrating an example embodiment of a system, such as a transmitter, that includes video encoding. [Figure 17] FIG. 2 is a block diagram illustrating another example embodiment of a video encoder. [Figure 18] FIG. 2 is a block diagram illustrating another example embodiment of a video encoder. [Figure 19] FIG. 2 is a block diagram illustrating another example embodiment of a video encoder. [Figure 20] 1 is a flow diagram illustrating an example embodiment of a process for encoding video. [Figure 21] 21 is a flow diagram illustrating an example of an embodiment of a feature of at least one other embodiment, such as the process shown in FIG. 20. [Figure 22] 10 is a graph illustrating features of an embodiment of an improved table. [Figure 23] 10 is a graph illustrating characteristics of another embodiment of a refinement table. [Figure 24]22 is a flow diagram illustrating an example embodiment of a variation of the process shown in FIG. 21. [Figure 25] 10 is a graph illustrating characteristics of another embodiment of a refinement table. [Figure 26] 10 is a graph illustrating characteristics of another embodiment of a refinement table. [Figure 27] 1 is a block diagram illustrating an example embodiment of a system such as a receiver that includes video decoding. [Figure 28] FIG. 2 is a block diagram illustrating another example embodiment of a video decoder. [Figure 29] FIG. 2 is a block diagram illustrating another example embodiment of a video decoder. [Figure 30] FIG. 2 is a block diagram illustrating another example embodiment of a video decoder. [Figure 31] 1 is a flow diagram illustrating an example embodiment of a process for decoding a video signal. [Figure 32] FIG. 2 is a block diagram illustrating another example embodiment of a video encoder. [Figure 33] FIG. 2 is a block diagram illustrating another example embodiment of a video decoder. DETAILED DESCRIPTION OF THE INVENTION
[0014] It should be understood that the drawings are intended to illustrate examples of various aspects and embodiments and are not necessarily the only possible configurations. Like reference symbols throughout the various drawings refer to the same or similar features.
[0015] Detailed Description The video compression and reconstruction process may introduce distortion. For example, distortion may be observed in the reconstructed video signal, particularly when the video signal is mapped before encoding to better utilize the distribution of codewords of samples in a video picture. Generally, at least one embodiment may provide flexibility and coding gain using video signal refinement, e.g., based on chroma components, allowing for reduced distortion and / or improved compression efficiency. At least one embodiment may include a refinement step as an in-loop filter or an out-of-loop filter to refine the reconstructed video signal after decoding. The refinement may be based on, e.g., a refinement table coded into the bitstream or signal during encoding. The decoder then applies signal correction (e.g., filtering) based on the refinement table. Among various examples of possible modes, one example of a mode in at least one embodiment is inter-component refinement of chroma components. Another example of a mode in at least one embodiment is intra-component refinement applied, e.g., to the luma component.
[0016] Applying a particular process, method, or technique intended to produce an improvement to an entire slice or to a given region may result in degradation of local portions of the slice or region to which the improvement is applied, even if the improvement is generally beneficial to the entire slice or region. Generally, at least one embodiment addresses this issue and also aims to improve the coding of the involved data, such as improvement tables. An improvement that addresses the entire slice or region may involve coding one table per component and applying that one table to all samples in the entire slice or region.
[0017] Generally, in at least one embodiment, a table is calculated at the encoder side based on providing an improvement indicated by a metric, for example, by reducing the rate-distortion cost (which is typically a weighted sum of distortion and coding cost) across the entire picture or region under consideration. The distortion may be, for example, the mean squared error between the improved reconstructed signal and the original signal. The resulting table is then coded into the bitstream or signal. If the rate-distortion cost across the considered slice or region decreases overall, this also means that there may be portions within the considered slice or region where the rate-distortion cost increases. Generally, at least one example embodiment may include an improvement mode that includes inter-component chroma improvement. Another example embodiment may include intra-component improvement applied, for example, to the luma component.
[0018] As used herein, improving data may refer to a change, adjustment, modification, revision, or adaptation of a process that may result in an improvement, such as an increase in the efficiency of coding and / or decoding information, such as video data, and / or other improvements, such as a reduction in cost, such as the rate-distortion cost described herein. For ease of explanation, one or more embodiments described herein may refer to a “cost” as being a “rate-distortion cost” that may be used to evaluate, determine, or obtain an effect, impact, or improvement associated with using an improved mode. For example, a cost associated with an improved mode may be evaluated (e.g., compared) to a cost without use of the improved mode, and a decision may be made to use an improved mode described herein based on an improvement, such as improved compression efficiency and / or reduced distortion, as indicated by a change in cost. While one type of cost may include a rate-distortion cost, references to rate-distortion cost herein are not limiting, and other techniques for evaluating other types of costs and / or the effect of an improved mode are contemplated.
[0019] Furthermore, in this description, the terms "reconstruct" and "decode" may be used interchangeably. Typically, but not necessarily, "reconstruct" is used on the encoder side, while "decode" is used on the decoder side. It should be noted that the terms "decode" or "reconstruct" may refer to partially "decoding" or "reconstructing" a bitstream or signal, e.g., a signal obtained after deblocking filtering but before SAO filtering, and the reconstructed samples may differ from the final decoded output used for display. Furthermore, the terms "image," "picture," and "frame" may be used interchangeably.
[0020] In general, at least one embodiment may include being able to activate block-by-block refinement and possibly selecting refinement parameters for application to the block from a set of possible refinement parameters previously coded in the stream. The term "block" may be replaced in variants by "coding unit" (CU), "coding tree unit" (CTU).
[0021] In general, one or more embodiments may relate to encoders and decoders, and one or more embodiments may relate to decoder specifications and semantics.
[0022] 1 shows an example embodiment of an encoder 100. Although variations of this encoder 100 are possible, encoder 100 is described below for clarity without describing all anticipated variations.
[0023] Before being encoded, the video sequence may be subjected to pre-encoding processing (101), for example, applying a color transformation to the input color picture (e.g., from RGB 4:4:4 to YCbCr 4:2:0) or remapping the input picture components to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with the pre-processing and may be added to the bitstream or signal.
[0024] Within encoder 100, a picture is coded by the encoder elements as described below. The picture to be coded is divided (102) into units of, for example, CUs and processed. Each unit is coded, for example, using intra mode or inter mode. When coding a unit in intra mode, intra mode performs intra prediction (160). In inter mode, motion estimation (175) and motion compensation (170) are performed. The encoder decides (105) whether to use intra mode or inter mode to code the unit and indicates the intra / inter decision, for example, with a prediction mode flag. A prediction residual is calculated, for example, by subtracting (110) the predicted block from the original image block.
[0025] The prediction residual is then transformed (125) and quantized (130). The quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded (145) to output a bitstream or signal. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., the residual is coded directly without applying the transform or quantization processes.
[0026] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (140) and inverse transformed (150) to decode the prediction residual. The decoded prediction residual is combined (155) with the predicted block to reconstruct an image block. An in-loop filter (165) is applied to the reconstructed picture, e.g., to perform deblocking / sample adaptive offset (SAO) filtering to reduce coding artifacts. The filtered image is stored in a reference picture buffer (180).
[0027] Figure 2 shows a block diagram of an example embodiment of a video decoder 200, in which a bitstream or signal is decoded by elements of the decoder as described below. The video decoder 200 generally performs a decoding path that is the inverse of the encoding path described in Figure 1. The encoder 100 also generally performs video decoding as part of encoding the video data.
[0028] Specifically, the decoder input includes a video bitstream or signal, such as may be generated by video encoder 100. The bitstream or signal is first entropy decoded (230) to obtain transform coefficients, motion vectors, and other coding information. Picture partition information indicates how the picture is partitioned. The decoder can then divide the picture according to the decoded picture partition information (235). The transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residual. The decoded prediction residual is combined with a predicted block (255) to reconstruct an image block. The predicted block can result from intra-prediction (260) or motion-compensated prediction (i.e., inter-prediction) (275) (270). An in-loop filter (265) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (280).
[0029] The decoded picture may further undergo post-decoding processing (285), such as an inverse color conversion (e.g., YCbCr 4:2:0 to RGB 4:4:4) or an inverse remapping that reverses the remapping process performed in pre-encoding processing (101). Post-decoding processing may use metadata derived in pre-encoding processing and signaled in the bitstream or signal.
[0030] Generally, in at least one embodiment, one or more features described herein may affect portions, sections, or modules of an encoder / decoder (codec), such as the in-loop filters (165 and 265) and entropy coding (145) and entropy decoding (230), according to various examples of the features described herein.
[0031] Figure 3 shows a variant 3000 of the video encoder 100 of Figure 1. Modules in Figure 3 that are identical to modules in Figure 1 are labeled with the same reference numbers and will not be described further. In Figure 3, the illustrated embodiment includes an in-loop refinement process. A step of calculating refinement parameters (step 3001) occurs after the in-loop filtering step (165). The refinement parameter calculation step takes as input the input video and the reconstructed video resulting from the in-loop filtering step. The refinement parameter calculation step (step 3001) calculates refinement parameters that are coded into the bitstream or signal by a modified entropy coder (3003). The refinement parameters are also used to apply (3002) the refinement step to the reconstructed video resulting from the in-loop filtering.
[0032] Figure 4 shows a variant 102 of the video encoder 100 of Figure 2. Modules in Figure 4 that are identical to modules in Figure 2 are labeled with the same reference numbers and will not be described further. In Figure 4, a bitstream or signal is decoded by an entropy decoder in step 4001. The entropy decoder produces a reconstructed video and additionally decoded refinement parameters. The refinement parameters are used in step 4002 to apply refinements of the reconstructed video to produce an output video.
[0033] As used herein, the following notations represent or correspond to the following: - B is the bit depth of the signal (e.g. 10 bits) - MaxVal is (2 B -1) is the maximum value of the signal - Sin#Y(p), Sin#cb(p), Sin#cr(p) correspond to the original (input) signals Y, cb, and cr at relative position p in the picture Sout#cb(p), Sout#cr(p) correspond to the (input) signals cb and cr resulting from the refinement at the relative position p in the picture. - Srec#Y(p) corresponds to the reconstructed luma signal at relative position p in the picture. This corresponds to the signal resulting from the in-loop filter. Srec#cb(p), Srec#cr(p) correspond to the reconstructed chroma signal to be improved at the relative position p in the picture. This corresponds to the signal resulting from the in-loop filter. - neutralVal corresponds to the default value, if the refinement value is equal to neutralVal the refinement will not modify the signal, a typical value for neutralVal is 64 - Rcb, Rcr correspond to chroma improvement parameters Rcb and Rcr can typically be defined as pivot points of a piecewise linear (PWL) model. For example, Rcb=[(Rcb#idx[pt], Rcb#val[pt]), pt=0 to Ncb-1], where Ncb is the number of pivot points.
[0034] The lookup tables LutRcb and LutRcr are derived from the chroma refinement parameters Rcb and Rcr. For example, LutRcb can be constructed by linear interpolation between each pair of points in Rcb as follows: pt=0 to Ncb-2 From idx=Rcb#idx[pt] to (Rcb#idx[pt+1]-1), LutRcb[idx]=Rcb#val[pt]+(Rcb#val[pt+1]-Rcb#val[pt])*(idx-Rcb#idx[pt]) / (Rcb#idx[pt+1]-Rcb#idx[pt])
[0035] One embodiment of inter-component chroma refinement works as follows (shown here for component cb): · Sout#cb(p)=offset+LutRcb[Srec#Y(p)] / neutralVal*(Srec#cb(p)-offset) However, offset is typically set to (MaxVal / 2), and p corresponds to the same relative position within the picture. Or it works like this: · Sout#cb(p)=(LutRcb[Srec#Y(p)]-neutralVal)+Srec#cb(p)
[0036] In another example embodiment, the refinement can be applied as an intra-component refinement of the luma component, which works as follows: · Sout#Y(p)=(LutRy[Srec#Y(p)] / neutralVal)*Srec#Y(p) Or it works like this: · Sout#Y(p)=LutRy[Srec#Y(p)]+Srec#Y(p)
[0037] One embodiment may include providing a bitstream or signal having one or more syntactic features as shown in the example of Figure 5. One aspect may include only the chroma components being refined. For example, if refinement_table_flag_cb is true, then for all pt from 0 to (refinement_table_cb_size-1), calculate the cb refinement table values (Rcb_idx[pt]), Rcb_val[pt]) as follows: Set Rcb#idx[pt] equal to (Rcb#idx[pt-1]+refinement#cb#idx[pt]) Set Rcb#val[pt] equal to (neutralVal+refinement#cb#value[pt]) A similar process is applied to the cr component.
[0038] An example of a refinement table is shown in Figure 6, where refinement#table#cb#size=6 (pt=0 to 5). For simplicity of notation, "Rcb" has been replaced with "R" in Figure 6.
[0039] In one embodiment, when block-based inter-component chroma refinement is activated, K pairs of cb, cr refinement tables (Rcb k , Rcr k), k=1 to K are enabled. If K is equal to 0, no refinement is applied within the slice. One or several block-based indicators are coded per block to indicate which of these table pairs to use, or to indicate whether refinement is not activated or not for this block.
[0040] In a non-limiting example embodiment, only one pair of tables is used (K=1).
[0041] In another example embodiment, a set of refinement tables is coded first in the slice header. An example of the syntax for decoding the cb, cr refinement table pair is shown in Figure 7. Figure 7 shows an example of slice header syntax using a table pair identifier: refinement#number#of#tables indicates the number of chroma refinement table pairs coded in the stream. This number shall be greater than or equal to 0. If it is equal to 0, it means that no chroma refinement is applied to the current slice. If refinement#number#of#tables is greater than 0, for any value of idx between 0 and (refinement#number#of#tables-1), at least refinement#table#flag#cb[idx] or refinement#table#flag#cr[idx] shall be equal to 1. Note: refinement#number#of#tables corresponds to the parameter K discussed above. refinement#table#flag#cb[idx] indicates whether a refinement table is coded for the cb component of the refinement-table pair with identifier value idx. If refinement#table#flag#cb[idx] is equal to 0, then refinement#table#flag#cr[idx] shall be equal to 1.
[0042] refinement_table_cb_size[idx] indicates the size of the cb refinement table of the refinement-table pair with identifier value idx. refinement#cb#idx[idx][i] denotes the index (first coordinate) of the i-th element in the cb refinement table of the refinement-table pair with identifier value idx. refinement#cb#value[idx][i] indicates the value (second coordinate) of the i-th element of the cb refinement table of the refinement table pair with identifier value idx. refinement#table#flag#cr[idx] indicates whether a refinement table is coded for the cr component of the refinement-table pair for identifier value idx. If refinement#table#flag#cr[idx] is equal to 0, then refinement#table#flag#cb[idx] shall be equal to 1. refinement_table_cr_size[idx] indicates the size of the cr refinement table for the refinement-table pair with identifier value idx. refinement#cr#idx[idx][i] denotes the index (first coordinate) of the i-th element in the cr refinement table of the refinement-table pair with identifier value idx. refinement#cr#value[idx][i] denotes the value (second coordinate) of the i-th element of the cr refinement table of the refinement-table pair with identifier value idx.
[0043] At the block (or CU or CTU) level, a block-based identifier tables_id is signaled, which corresponds to the identifier of the cb, cr refinement table pair that should be applied to the current block. An example of the syntax for decoding block refinement information is shown in Figure 8, which shows an example of a block-level syntax using table pair identifiers. In Figure 8: The syntax element refinement_activation_flag indicates whether the current block uses refinement. If the current block uses refinement (refinement#activation#flag is equal to 1), decode the syntax element tables#id. Refine the current block using the refinement table corresponding to the identifier tables#id: i=0 to (refinement#table#cb#size[tables#id]-1) 〇 refinement#cb#idx[tables#id][i] 〇 refinement#cb#value[tables#id][i] i=0 to (refinement#table#cr#size[tables#id]-1) 〇 refinement#cr#idx[tables#id][i] 〇 refinement#cr#value[tables#id][i]
[0044] A corresponding block diagram illustrating an embodiment for deriving refinement information for the current block is shown in Figure 9. In Figure 9, the syntax element refinement_activation_flag is decoded in 601. In 602, it is checked whether refinement_activation_flag is equal to 0. If refinement_activation_flag is equal to 0, no refinement is applied (603). If refinement_activation_flag is not equal to 0, tables_id is decoded in 604. In 605, the refinement is applied using the table pair indicated by the identifier tables_id.
[0045] In one embodiment, the syntax element refinement_activation_flag is not coded, but a specific value of tables_id indicates that the block is not refined. For example, tables_id equal to 0 indicates that the current block does not use refinement. The test "refinement_activation_flag==0?" is replaced by "tables_id==0?".
[0046] In another embodiment, an aspect includes coding relative to neighboring refinement information, for example, a block-based indicator is created in the syntax element that indicates whether the refinement information for the current block is replicated from a neighboring block. - A parameter indicating whether the block used refinement (here named refinedBlock, which corresponds to the syntax element named refinement#activation#flag) - The refinement table used to refine samples of the block if refinedBlock indicates that the block is refined. It is made with. An example of the syntax for decoding block refinement information is shown in Figure 10: tables_from_left_flag indicates whether refinement information is replicated from the block to the left of the current block. tables_from_up_flag indicates whether refinement information is replicated from the block above the current block. refinement_activation_flag indicates whether refinement is activated for the current block.
[0047] The corresponding block diagram for deriving the refinement information of the current block is shown in Figure 11. In Figure 11, (501) decodes the syntax element tables#from#left#flag. 502 checks whether tables#from#left#flag is equal to 0. If tables#from#left#flag is not equal to 0, the refinement information of the current block is copied from the block to the left of the current block (504). Otherwise, (503) decodes the syntax element tables#from#up#flag. 505 checks whether tables#from#up#flag is equal to 0. If tables#from#up#flag is not equal to 0, the refinement information of the current block is copied from the block above the current block (507). Otherwise, (506) decodes the syntax element refinement#activation#flag. The refinedBlock of the current block is set to refinement#activation#flag. 508 checks whether refinement#activation#flag is equal to 0. If refinement_activation_flag is equal to 1, the refinement table is decoded (510). Otherwise, no specific process is applied (510). The refinement is applied in 511. In this step, if refinedBlock is equal to 0, the block is not processed by the refinement process. If refinedBlock is equal to 1, the refinement table derived for the current block is used to refine the block.
[0048] In one embodiment, refinement_activation_flag is not explicitly coded or decoded, but is implicitly inferred contextually. For example, the context relates to the SAO loop filtering process. In this embodiment, refinement_activation_flag is deduced from the SAO parameters. In one embodiment, if SAO is disabled for a block, refinement_activation_flag is set equal to 0. If SAO is enabled for a block, refinement_activation_flag is set equal to 1. In another embodiment, if SAO band offset is disabled for a block, refinement_activation_flag is set equal to 0. If SAO band offset is enabled for a block, refinement_activation_flag is set equal to 1. Similar rules can be applied for other SAO types, such as edge offset.
[0049] In one embodiment, the refinement process is applied together with another loop filter process, such as SAO.
[0050] In general, one embodiment may include reducing the coding cost of the table by avoiding redundant neutral values. In one embodiment, consider that the chroma refinement tables have the same size (Ncb=Ncr=N) and share the same index: Rcb#idx[pt]=Rcr#idx[pt], pt=0 to (N-1) It is observed that in most cases, when the refinement parameter of one component (say cb) is equal to the neutral value neutralVal, the corresponding refinement parameter of the other chroma component will often be close to the neutral value.
[0051] This is illustrated by the following example, which is taken from processing the first picture of the SDR content BasketballDrive#1 920x1080, which is coded using the VTM codec with a QP of 37. The neutralValue is equal to 64, and the cases where the refinement value is close to 64 are shown in bold font.
number
[0052] In the encoder, the input values are an improvement parameter Rp#c1 of a first chroma component (e.g., cb component) and a corresponding improvement parameter Rp#c2 of a second chroma component (specified with respect to the same relative index in the improvement parameter table). The first step is to encode the improvement parameter Rp#c1 of the first chroma component (step 701). The second step is to check a condition related to the value of Rp#c1 (step 702). If the condition is false, the value of the corresponding improvement parameter Rp#c2 of the second chroma component is encoded (step 704). If the condition is true, the value of Rp#c2 is inferred (step 703).
[0053] In the decoder, the first step is to decode the refinement parameter Rp#c1 of the first chroma component (step 711). The second step involves checking a condition related to the value of Rp#c1 (step 712). If the condition is false, the value of the corresponding refinement parameter Rp#c2 of the second chroma component is decoded (step 714). If the condition is true, the value of Rp#c2 is inferred (step 713).
[0054] In one embodiment, the condition tested may include checking whether the value of the refinement parameter for the first chroma component is equal to neutralValue: Rp#c1==neutralValue?
[0055] In another embodiment, the condition tested may include checking whether the value of the refinement parameter for the first chroma component is greater than or equal to (neutralValue-thresh1) or less than or equal to (neutralValue-thresh2), where thresh1 and thresh2 are default values (typically equal to 1): Rp#c1>=(neutralValue-thresh1) AND Rp#c1<=(neutralValue+thresh2)?
[0056] In one embodiment, the inferred value of Rp#c2 is neutralValue.
[0057] In one embodiment, the inferred value of Rp#c2 is Rp#c1.
[0058] In one embodiment, the inferred value of Rp#c2 is (neutralValue-Rp#c1).
[0059] An example of the modified syntax with the condition Rp#c1(refinement#cb#value[i] in table) == neutralValue? is shown in Figure 13. If this condition is true, Rp#c2(refinement#cr#value[i] in table) is inferred to neutralValue. Otherwise, this element is explicitly signaled.
[0060] In general, an embodiment may include controlling the size of the refinement table. For example, an embodiment may include at least one of the following to reduce the coding cost of the refinement table: The table size of cb and cr is the same, fixed and determined by default, here named N. · The indices in the tables Rcb#idx[pt] and Rcr#idx[pt] correspond to the equidistant points. A limited predefined number Nref of elements are coded per refinement table. In this case, the indication pt#init of the first table point at which the elements of the table are coded is given. The refinement value to be coded is therefore 〇 Rcb#idx[pt#init] to Rcb#idx[pt#init+Nref-1] ○ Rcr#idx[pt#init] to Rcr#idx[pt#init+Nref-1]
[0061] Other (uncoded) values are set to a neutral value (NeutralVal). ○ Rcr#idx[pt]=neutralVal and Rcb#idx[pt]=neutralVal, pt=0 to (pt#init-1) and (pt#init+Nref) to (N-1) In one embodiment, the number of values per refinement table, Nref, is equal to four.
[0062] In one embodiment, the encoder may determine block-based refinement activations and derive one or more refinement tables according to one or more embodiments described herein based on, for example, the actions listed below and illustrated in the corresponding block diagrams shown in Figure 14. In one implementation, the listed actions may be performed iteratively. 1. Initialize the map blkTablesId[blkId] to -1, where blkId=0 to (Nblk-1), where Nblk is the number of blocks in the slice (step 802). 2. The table identifier parameter tablesId is initialized to 0 (step 803). 3. Derive refinement table with identifier tablesId from samples in block blkId where blkTablesId[blkId]=-1 (step 804). a. The result is the cb and cr tables: i. (refinement#cb#idx[tablesId][i], refinement#cb#value[tablesId][i]), where i ranges from 0 to (refinement#table#cb#size[tablesId] - 1), and (refinement#cr#idx[tablesId][i], refinement#cr#value[tablesId][i]), where i ranges from 0 to (refinement#table#cr#size[tablesId] - 1) b. For the embodiments for performing the described calculations, see the following. For each block blkId where blkTablesId[blkId] = -1, calculate the chroma rate distortion cost without improvement (RD#without[blkId]) and the chroma rate distortion cost with improvement (RD#with[blkId]) using the table calculated in step 3 (step 805). a. Calculate the chroma rate distortion cost as follows: distortion(cb) + distortion(cr) + L * cost(refinement#table#cb) + L * cost(refinement#table#cr) + L * cost(activation map) where L is the well-known "lambda" parameter in the derivation of the rate distortion cost. 5. For all block blkIds where blkTablesId[blkId] = -1, update the block-based improvement activation as follows (step 806): a. If (RD#with[blkId] < RD#without[blkId]) holds, then blkTablesId[blkId] = tablesId 6. If there are still block blkIds such as blkTablesId[blkId] = -1 and tablesId < MaxNbTables (step 807), a. Increment tablesId by 1 and proceed to step 3. b. Otherwise, stop (step 809).
[0063] MaxNbTables is a default parameter that specifies the maximum number of enabled pairs in the refinement table (step 808).
[0064] Regarding aspect 3b above, in one embodiment, refinement metadata can be derived as described below. The refinement metadata represents a correction function, denoted R(), that is applied to individual samples of a color component (e.g., the luma component Y, or the chroma components Cb / Cr). One aspect of the refinement metadata is to minimize the rate-distortion cost between the reconstructed picture and the original picture. This minimization can be performed over a given region A, e.g., the whole picture, a slice, a tile, or a CTU. The refinement metadata is denoted R. In a preferred implementation, a piecewise linear model defined by N pairs (R#idx[k], R#val[k]), k=0 to N-1, is used for R.
[0065] The process begins by initializing the refinement metadata R with a neutral value (step 300). A neutral value is one where the refinement does not change the signal. This is explained further below. R is used to calculate an initial rate-distortion cost, initRD. The derivation of the rate-distortion cost using the function R is explained below. The parameter bestRD is initialized to initRD. The step of optimizing the refinement metadata R consists of the following steps: Run a loop over the indices pt of successive pivot points R of the PWL model. Initialize the parameters bestVal and initVal to R#val[pt]. Run a loop over various values of R#val[pt]. The loop runs from a value of (initVal-Val0) to a value of (initVal+Val1), where Val0 and Val1 are default parameters. Typical values are shown below: Calculate the rate-distortion cost, curRD, using the revised R (along with the revised R#val[pt]). A comparison of curRD with bestRD is shown. If curRD is lower than bestRD, set bestRD to curRD and set bestValue to R#val[pt]. Then check for loop termination on the value of R#val[pt]. If the loop is complete, set R#val[pt] to bestValue. Then check the loop on the value of pt. If the loop is completed, the resulting metadata of the optimization process is R. The step of optimizing the refinement metadata can be repeated several times.
[0066] The following notation will now be used to describe an embodiment of the improved process: - B is the bit depth of the signal (e.g. 10 bits) - MaxVal is (2 B -1) is the maximum value of the signal - Sin(p) is the original (input) signal at position p in the picture Sout(p) is the signal resulting from the refinement at position p in the picture Srec(p) is the signal to be improved. It corresponds to the signal resulting from the inverse mapping in the first variant or the signal resulting from the in-loop filter in the second variant. - NeutralVal corresponds to the default value, if the refinement value is equal to NeutralVal the refinement will not modify the signal, a typical value for NeutralVal is 128
[0067] A lookup table LutR can be constructed from the refinement metadata R as follows: pt=0 to N-2 From idx=R#idx[pt] to (R#idx[pt+1]-1), LutR[idx]=R#val[pt]+(R#val[pt+1]-R#val[pt])*(idx-R#idx[pt])*(R#idx[pt+1]-R#idx[pt])
[0068] Generally, at least two examples of refinement modes are defined, referred to herein as "Mode 1" or "intra-component refinement mode," and "Mode 2" or "inter-component refinement mode": Mode 1 or intra-component refinement mode performs the refinement independently of the other components. Preferably, this mode is applied to the luma component and is based on the following formula: 〇 Sout(p)=LutR[Srec(p)] / NeutralVal*Srec(p) In mode 2 or inter-component refinement mode, refinement is performed on one component depending on another component. Preferably, this mode is applied to the chroma components with a dependency on the luma component (hereafter denoted Srec#Y(p)). Note that Srec#Y(p) can be the signal resulting from an in-loop filter step or from an inverse mapping in a first variant. Srec#Y(p) can also be a filtered version of the other (Y) component. This process is based on the following formula: 〇 Sout(p)=offset+LutR[Srec#Y(p)] / NeutralVal*(Srec(p)-offset) However, offset is typically set to (MaxVal / 2), and p corresponds to the same relative position within the picture. Rounding and clipping between the minimum and maximum signal values (typically between 0 and 1023 for a 10-bit signal) is finally applied to the refined value Sout(p). Here the refinement is applied as a multiplicative operator mode. The refinement can also use an additive refinement operator. In this case mode 1 is applied as follows: 〇 Sout(p)=LutR[Srec(p)] / NeutralVal+Srec(p) Mode 2 is applied as follows: 〇 Sout(p)=LutR[Srec#Y(p)] / NeutralVal+Srec(p) For the chroma components, refinement mode 2 with multiplicative operator mode is preferred.
[0069] Then, in one embodiment, the rate-distortion cost of the refinement metadata R can be determined as follows, where the following notation is used: A is the picture area to refine, the refinement can be done for example on the whole picture, a slice, a tile, a CTU. - Dist(x,y) is the distortion between sample value x and sample value y. Typically, the distortion is the squared error (xy) 2 is - Cost(R) is the coding cost for coding the refinement metadata R - L is the lambda coefficient associated with the region. L is typically 2 (QP / 6) , and QP represents the quantization parameter applied to the region. The rate-distortion cost RDcost using the refinement metadata R is defined as:
number
[0070] Generally, an embodiment provides a coding gain, i.e., an increase in quality for the same bit rate or a decrease in bit rate for the same quality. The refinement gain due to block-based activation is shown in the table below. In the non-limiting example shown below, only one table per component is allowed (K or equivalently, refinement_number_of_tables equals 1).
[0071] Performance is shown for 5 HDR HD 10-bit content, 5 SDR HD 8-bit content, and 2 SDR HD 10-bit content. The codec used is VTM (Versatile Video Coding (VVC) Test Model) with QTBT (quadtree and binary tree block partitioning structure) activated. All sequences are coded with an internal bit depth of 10 bits. The table below shows the coding results for a sequence of 17 frames in random access configuration. neutraVal is set to 64 and the table size is set to 17. Improved metadata coding is only applied for temporal level 0.
[0072] The following five tables show the BD rate gains as follows: Table 1 when applying the refinement to all blocks ("Full Slice Refinement"), Tables 2 to 4 when enabling activation of refinement per block of size 128x128 ("blk128 Refinement"), refinement per block of size 64x64 ("blk64 Refinement"), and refinement per block of size 32x32 ("blk32 Refinement"), and Table 5 when selecting the best configuration per content. These results show the benefit or improvement of using block-based refinement activation.
[0073] [Table 1]
[0074] [Table 2]
[0075] [Table 3]
[0076] [Table 4]
[0077] [Table 5]
[0078] This specification describes various example embodiments, features, models, techniques, etc. Many such examples are described with specificity, and often in a manner that may be considered limiting, at least to illustrate their individual characteristics, but for clarity of description only and not to limit the application or scope of this specification. Indeed, various example embodiments, features, etc. described herein can be combined and interchanged in various ways to produce further example embodiments.
[0079] Generally, the example embodiments described and contemplated herein can be implemented in many different forms. While Figures 1 and 2 above and Figure 15 below illustrate some embodiments, other embodiments are contemplated, and the discussion of Figures 1, 2, and 15 is not intended to limit the scope of implementations. At least one embodiment provides an example generally related to encoding and / or decoding video, and at least one other embodiment generally related to transmitting a generated or encoded bitstream or signal. These and other embodiments can be implemented as a method, an apparatus, a computer-readable storage medium storing instructions for encoding or decoding video data according to any of the described methods, and / or a computer-readable storage medium storing a bitstream or signal generated according to any of the described methods.
[0080] In this application, the terms "reconstruct" and "decode" may be used interchangeably, the terms "pixel" and "sample" may be used interchangeably, and the terms "image," "picture," and "frame" may be used interchangeably. Typically, but not necessarily, the term "reconstruct" is used on the encoder side, while "decode" is used on the decoder side.
[0081] This disclosure uses the terms HDR (high dynamic range) and SDR (standard dynamic range). These terms often convey specific values of dynamic range to those skilled in the art. However, additional embodiments are contemplated in which reference to HDR is understood to mean "higher dynamic range" and reference to SDR is understood to mean "lower dynamic range." Such additional embodiments are not constrained by any specific values of dynamic range that may often be associated with the terms "high dynamic range" and "standard dynamic range."
[0082] Various methods are described herein, each of which includes one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for the method to operate properly, the order and / or use of specific steps and / or actions can be modified or combined.
[0083] Various methods and other aspects described herein can be used to modify modules, such as the intra-prediction, entropy coding, and / or decoding modules (160, 360, 145, 330) of video encoder 100 and decoder 200 shown in Figures 1 and 2. Furthermore, aspects herein are not limited to VVC or HEVC, but can be applied, for example, to other standards and recommendations, existing or developed in the future, and extensions of any such standards and recommendations (including VVC and HEVC). Unless otherwise specified or technically excluded, aspects described herein can be used individually or in combination.
[0084] For example, various numerical values are used herein, and the specific values are for illustrative purposes only, and the described aspects are not limited to those specific values.
[0085] FIG. 15 illustrates a block diagram of an example system in which various aspects and embodiments can be implemented. System 1000 can be implemented as a device including various components described below and configured to perform one or more of the aspects described herein. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected consumer electronics, and servers. Elements of system 1000 can be implemented singly or in combination within a single integrated circuit, multiple ICs, and / or separate components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 1000 are distributed across multiple ICs and / or separate components. In various embodiments, system 1000 is communicatively coupled to other similar systems or other electronic devices, for example, via a communication bus or by dedicated input and / or output ports. In various embodiments, system 1000 is configured to implement one or more of the aspects described herein.
[0086] The system 1000 includes at least one processor 1010 configured to execute instructions loaded therein, for example, to implement various aspects described herein. The processor 1010 may include embedded memory, input / output interfaces, and various other circuits known in the art. The system 1000 includes at least one memory 1020 (e.g., a volatile memory device and / or a non-volatile memory device). The system 1000 includes storage 1040, which may include non-volatile memory and / or volatile memory including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drives, and / or optical disk drives. The storage 1040 may include, by way of non-limiting example, internal storage, additional storage, and / or network-accessible storage.
[0087] System 1000 includes an encoder / decoder module 1030 configured to process data to provide, for example, encoded video or decoded video, which may include its own processor and memory. The encoder / decoder module 1030 represents a module that may be included within a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding module and a decoding module. Additionally, the encoder / decoder module 1030 may be implemented as a separate element of system 1000 or may be incorporated within processor 1010 as a combination of hardware and software, as is known to those skilled in the art.
[0088] Program code loaded onto the processor 1010 or the encoder / decoder 1030 to perform various aspects described herein may be stored in storage 1040 and then loaded onto the memory 1020 for execution by the processor 1010. According to various embodiments, one or more of the processor 1010, memory 1020, storage 1040, and encoder / decoder module 1030 may store one or more of various items during execution of the processes described herein. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams or signals, matrices, variables, and intermediate or final results of equations, formulas, operations, and processing of arithmetic logic.
[0089] In some embodiments, memory internal to the processor 1010 and / or encoder / decoder module 1030 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing unit (e.g., the processing unit may be the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory may be memory 1020 and / or storage 1040, such as dynamic volatile memory and / or non-volatile flash memory. In some embodiments, external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, high-speed external dynamic volatile memory, such as RAM, is used as working memory for video coding and decoding operations, such as MPEG-2, HEVC, or Versatile Video Coding (VVC).
[0090] Input to the elements of system 1000 may be provided by various input devices shown in block 1130. Such input devices include, but are not limited to, (i) an RF section that receives RF signals transmitted wirelessly, for example, by a broadcaster, (ii) a composite input, (iii) a USB input, and / or (iv) an HDMI input.
[0091] In various embodiments, the input devices of block 1130 have associated individual input processing elements known in the art. For example, the RF section may be associated with elements for (i) selecting a desired frequency (also referred to as selecting a signal or bandlimiting a signal to a frequency band), (ii) downconverting the selected signal, (iii) bandlimiting again to a narrower frequency band to select a signal frequency band, which may be referred to as a channel in certain embodiments (for example), (iv) demodulating the downconverted and bandlimited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF section of various embodiments includes one or more elements for performing these functions, such as a frequency selector, a signal selector, a bandlimiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner for performing various of these functions, including, for example, downconverting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency close to baseband) or to baseband. In some set-top box embodiments, the RF section and its associated input processing elements receive RF signals transmitted over a wired (e.g., cable) medium and perform frequency selection by filtering, downconverting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of those elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.
[0092] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 1000 to other electronic devices across the USB and / or HDMI connections. It should be understood that various aspects of input processing, e.g., Reed-Solomon error correction, can be implemented, for example, in a separate input processing IC or in processor 1010. Similarly, aspects of USB or HDMI interface processing can be implemented in a separate interface IC or in processor 1010. To process the data stream for presentation on an output device, the modulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 1010 and encoder / decoder 1030 operating in combination with memory and storage elements.
[0093] The various elements of system 1000 may be provided within a unitary housing in which the various elements are interconnected and may transmit data therebetween using a suitable connection arrangement 1140, such as an internal bus known in the art, including an I2C bus, wiring, and a printed circuit board.
[0094] System 1000 includes a communication interface 1050 that enables communication with other devices over a communication channel 1060. Communication interface 1050 may include, but is not limited to, a transceiver configured to transmit and receive data over communication channel 1060. Communication interface 1050 may include, but is not limited to, a modem or a network card, and communication channel 1060 may be implemented in a wired and / or wireless medium, for example.
[0095] In various embodiments, data is streamed to system 1000 using a Wi-Fi network, such as IEEE 802.11. The Wi-Fi signal in these embodiments is received over communication channel 1060 and communication interface 1050, which are adapted for Wi-Fi communication. Communication channel 1060 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to enable streaming applications and other over-the-top communications. Other embodiments provide streamed data to system 1000 using a set-top box that delivers data over the HDMI connection of input block 1130. Still other embodiments provide streamed data to system 1000 using the RF connection of input block 1130.
[0096] System 1000 can provide output signals to various output devices, including display 1100, speakers 1110, and other peripherals 1120. In various example embodiments, other peripherals 1120 include one or more of a stand-alone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functionality based on the output of system 1000. In various embodiments, control signals are communicated between system 1000 and display 1100, speakers 1110, or other peripherals 1120 using signaling such as AV.Link, CEC, or other communication protocols that allow inter-device control with or without user intervention. Output devices may be communicatively coupled to system 1000 via dedicated connections through individual interfaces 1070, 1080, and 1090. Alternatively, output devices may be connected to system 1000 using communication channel 1060 via communication interface 1050. The display 1100 and speakers 1110 may be integrated into a single unit along with the other components of the system 1000 in an electronic device, such as a television. In various embodiments, the display interface 1070 includes a display driver, such as a timing controller (T Con) chip.
[0097] For example, if the RF portion of input 1130 is part of a separate set-top box, display 1100 and speakers 1110 can alternatively be separate from one or more of the other components. In various embodiments in which display 1100 and speakers 1110 are external components, the output signal can be provided by a dedicated output connection including, for example, an HDMI port, a USB port, or a COMP output.
[0098] The embodiments may be performed by computer software implemented by the processor 1010, by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. The memory 1020 may be of any type suitable for the technical environment and may be implemented using any suitable data storage technology, including, by way of non-limiting examples, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. The processor 1010 may be of any type suitable for the technical environment and may include, by way of non-limiting examples, one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.
[0099] Reference is now made to FIGS. 16 through 30, which generally illustrate various other example embodiments suitable for implementing one or more aspects described herein.
[0100] FIG. 16 illustrates an example of an architecture of a transmitter 1000 configured to encode pictures into a bitstream, according to a particular and non-limiting embodiment.
[0101] The transmitter 1000 includes one or more processors 1005, which may include, for example, a CPU, a GPU, and / or a DSP (an English acronym for digital signal processor), along with built-in memory 1030 (e.g., RAM, ROM, and / or EPROM). The transmitter 1000 includes one or more communication interfaces 1010 (e.g., a keyboard, a mouse, a touchpad, a webcam), each adapted to display output information and / or allow a user to input commands and / or data, and a power source 1020, which may be external to the transmitter 1000. The transmitter 1000 may also include one or more network interfaces (not shown). The encoder module 1040 represents a module that may be included within a device to perform coding functions. Additionally, the encoder module 1040 may be implemented as a separate element of the transmitter 1000 or may be incorporated within the processor 1005 as a combination of hardware and software, as known to those skilled in the art.
[0102] According to different embodiments, the picture can be obtained from sources including, but not limited to: - Local memory, e.g. video memory, RAM, flash memory, hard disk, - a storage interface, e.g. an interface to a mass storage, ROM, optical disk, or magnetic support; a communication interface, such as a wired interface (e.g. a bus interface, a wide area network interface, a local area network interface) or a wireless interface (such as an IEEE 802.11 interface or a Bluetooth interface), and - Picture capture circuitry (e.g. sensors such as CCD (i.e. charge coupled device) or CMOS (i.e. complementary metal oxide semiconductor)) It could be.
[0103] According to different embodiments, the bitstream can be transmitted to a destination. By way of example, the bitstream is stored in a remote or local memory, such as a video memory, or a RAM, or a hard disk. In a variant, the bitstream is transmitted to a storage interface, such as an interface to a mass storage, a ROM, a flash memory, an optical disk, or a magnetic support, and / or transmitted over a communication interface, such as an interface to a point-to-point link, a communication bus, a point-to-multipoint link, or a broadcast network.
[0104] According to a non-limiting example embodiment, the transmitter 1000 further includes a computer program stored in memory 1030. The computer program includes instructions that, when executed by the transmitter 1000, specifically by the processor 1005, enable the transmitter 1000 to perform the encoding method described with respect to FIG. 20 . According to a variant, the computer program is stored external to the transmitter 1000 on a non-transitory digital data support, for example, on an external storage medium such as a HDD, a CD-ROM, a DVD, a read-only and / or DVD drive, and / or a DVD read / write drive, all known in the art. The transmitter 1000 therefore includes a mechanism for reading the computer program. Furthermore, the transmitter 1000 can access one or more USB-type storage devices (e.g., “memory sticks”) via a corresponding Universal Serial Bus (USB) port (not shown).
[0105] According to one or more non-limiting example embodiments, the transmitter 1000 may include, but is not limited to: - mobile devices, - communication devices, - gaming consoles, - a tablet (or tablet computer), - laptop, - still image camera, - video camera, - coding chips or coding devices / equipment, - a still image server, and - Video servers (e.g. broadcast servers, video-on-demand servers, or web servers) It could be.
[0106] Figure 17 shows an example of a video encoder 100, for example an encoder of the HEVC type, adapted to perform the encoding method of Figure 20. The encoder 100 is an example of a transmitter 1000 or a part of such a transmitter 1000.
[0107] For coding purposes, a picture is typically divided into basic coding units, such as coding tree units (CTUs) in HEVC or macroblock units in H.264. A set of possibly consecutive basic coding units is grouped into a slice. A basic coding unit includes basic coding blocks for all color components. In HEVC, the smallest coding tree block (CTB) has a size of 16x16, which corresponds to the size of a macroblock used in previous video coding standards. While the terms CTU and CTB are used herein to describe encoding / decoding methods and devices, it will be understood that these methods and devices should not be limited by those specific terms, which may be named differently (e.g., macroblocks) in other standards such as H.264.
[0108] In HEVC coding, a picture is divided into square CTUs with configurable sizes, typically 64x64, 128x128, or 256x256. A CTU is the root of a quadtree partition into four square coded units (CUs) of equal size, i.e., half the width and height of the parent block. A quadtree is a tree in which a parent node can be divided into four child nodes, and each child node can be the parent node of another division into four child nodes. In HEVC, coded blocks (CBs) are divided into one or more predictive blocks (PBs), which form the root of a quadtree partition into transform blocks (TBs). Corresponding to coded blocks, predictive blocks, and transform blocks, coded units (CUs) contain predictive units (PUs) and a set of tree-structured transform units (TUs), where a PU contains prediction information for all color components and a TU contains a residual coding syntax structure for each color component. The sizes of the CBs, PBs, and TBs for the luma component apply to the corresponding CUs, PUs, and TUs.
[0109] In more recent coding systems, a CTU is the root of a coding tree division into coded units (CUs). A coding tree is a tree in which a parent node (usually corresponding to a CU) can be divided into child nodes (e.g., two, three, or four child nodes), and each child node can be the parent node of another division into child nodes. In addition to the quadtree partitioning mode, new partitioning modes (binary tree symmetric partitioning mode, binary tree asymmetric partitioning mode, and ternary tree partitioning mode) are also defined to increase the total number of possible partitioning modes. A coding tree has a unique root node, e.g., a CTU. The leaves of the coding tree are the terminal nodes of the tree. Each node of the coding tree represents a CU that can be further divided into smaller CUs, also called sub-CUs or, more generally, sub-blocks. Once the division of a CTU into CUs has been determined, the CUs corresponding to the leaves of the coding tree are coded. The division of a CTU into CUs and the coding parameters used to code each CU (corresponding to the leaves of the coding tree) can be determined at the encoder side by a rate-distortion optimization procedure. There is no division of CB into PB and TB, i.e. a CU is made up of a single PU and a single TU.
[0110] Hereinafter, the term "block" or "picture block" may be used to refer to any one of CTU, CU, PU, TU, CB, PB, and TB. In addition, the term "block" or "picture block" may be used to refer to macroblocks, partitions, and sub-blocks defined within H.264 / AVC or other video coding standards, and more broadly to refer to arrays of samples of various sizes.
[0111] Returning to Figure 17, in an example embodiment of encoder 100, the elements of the encoder encode a picture as described below. The picture to be encoded is processed in units of CUs. Each CU is encoded using intra mode or inter mode. When a CU is encoded in intra mode, intra mode performs intra prediction (160). In inter mode, motion estimation (175) and motion compensation (170) are performed. The encoder decides (105) whether to use intra mode or inter mode to encode the CU, and indicates the intra / inter decision with a prediction mode flag. A residual is calculated by subtracting (110) a predicted sample block (also known as a predictor) from the original picture block.
[0112] Intra-mode CUs are predicted from reconstructed neighboring samples within the same slice, for example. HEVC provides a set of 35 intra-prediction modes, including DC, planar, and 33 angular prediction modes. Inter-mode CUs are predicted from reconstructed samples of reference pictures stored in a reference picture buffer (180).
[0113] The residual is transformed (125) and quantized (130). The quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded (145) to output a bitstream. The encoder can also skip the transform or bypass both the transform and quantization, i.e., the residual is coded directly without applying the transform or quantization processes.
[0114] Entropy coding can be, for example, context-adaptive binary arithmetic coding (CABAC), context-adaptive variable length coding (CAVLC), Huffman, arithmetic, exponential-Golomb, etc. CABAC is an entropy coding method first introduced in H.264 and also used in HEVC. CABAC includes binarization, context modeling, and binary arithmetic coding. Binarization maps syntax elements to binary symbols (bins). Context modeling determines the probability of each regularly coded bin (i.e., non-bypass) based on a certain context. Finally, binary arithmetic coding compresses the bins into bits according to the determined probabilities.
[0115] The encoder includes a decoding loop, which decodes the coded block and provides a reference for further prediction. The quantized transform coefficients are dequantized (140) and inverse transformed (150) to decode the residual. A picture block is reconstructed by combining (155) the decoded residual with the predicted sample block. Optionally, an in-loop filter (165) is applied to the reconstructed picture to reduce coding artifacts, for example, by performing a deblocking filter (DBF) / sample adaptive offset (SAO) / adaptive loop filtering (ALF). The filtered picture is stored in a reference picture buffer (180) and can be used as a reference for other pictures. In this embodiment, refinement data is determined (190) from the filtered reconstructed picture, i.e., the output of the in-loop filter, and its original version. In a first variant, refinement data is determined (190) from the reconstructed picture before in-loop filtering and its original version. In this first variant, refinement is applied before in-loop filtering. In a second variant, refinement data is determined (190) from a partially filtered reconstructed picture, e.g., after deblocking filtering but before SAO, and its original version. In this second variant, the refinement is applied after the partial filtering, e.g., immediately after, e.g., after deblocking filtering but before SAO. The refinement data represents a correction function, denoted R(), that is applied to individual samples of a color component (e.g., luma component Y or chroma components Cb / Cr, or color components R, G, or B). The refinement data is then entropy coded into the bitstream. In this embodiment, the refinement process is outside the decoding loop. Therefore, the refinement process is applied in the decoder only as a post-processing step.
[0116] Figure 18 shows a variant 101 of the video encoder 100 of Figure 17. Modules in Figure 18 that are identical to those in Figure 17 are labeled with the same reference numbers and will not be described further. A picture can be mapped (105) before encoding by the video encoder 101. Such mapping can be used to better exploit the distribution of codeword values of the samples of the picture. As shown in Figure 18, mapping is typically applied to the original (input) samples before core encoding. Typically, a static mapping function is used, i.e., the same function for all content, to limit complexity. A mapping function fmap(), possibly modeled by a 1D lookup table LUTmap[x] (x is a value), is applied directly to the input signal x as follows: y=fmap(x) or y=LUTmap[x] where x is the input signal (for example, 0 to 1023 for a 10-bit signal) and y is the mapped signal.
[0117] In a variant, the mapping function for one component depends on another component (inter-component mapping function). For example, a chroma component c is mapped depending on a luma component y at the same relative position in the picture. The chroma component c is mapped as follows: c=offset+fmap(y)*(c-offset) or c=offset+LUTmap[y]*(c-offset) However, the offset is typically the center value of the chroma signal (e.g., 512 for a 10-bit chroma signal). This parameter can also be a dynamic parameter coded into the stream that can result in improved compression gain.
[0118] The mapping function can be predefined or can be signaled in the bitstream using, for example, a piecewise linear model, a scaling table, or a delta QP (dQP) table.
[0119] The filtered reconstructed picture, i.e., the output of the in-loop filter, is inversely mapped (185). Inverse mapping (185) implements the reverse process of mapping (105). Refinement data is determined (190) from the inversely mapped filtered reconstructed picture and its original version. The refinement data is then entropy coded into the bitstream. In this embodiment, the refinement process is outside the decoding loop; therefore, it is applied in the decoder only as a post-process.
[0120] Figure 19 shows a variation 102 of the video encoder 100 of Figure 17. Modules in Figure 19 that are identical to modules in Figure 17 are labeled with the same reference numbers and will not be described further. The mapping module 105 is optional. It determines (190) refinement data from the filtered reconstructed picture, i.e., the output of the in-loop filter and its original version if mapping is not applied, or from its mapped original version if mapping is applied. The refinement data is then entropy coded (145) into the bitstream. The refinement data is also used to refine (182) the filtered reconstructed picture. The refined picture, rather than the filtered reconstructed picture, is stored in the reference picture buffer (180). In this embodiment, the refinement process is an in-loop process, i.e., part of the decoding loop. Thus, the refinement process is applied in both the encoder and the decoder's decoding loop. Modules 182 and 190 can be inserted in various positions. The refinement module 182 can be inserted before the in-loop filter or between the in-loop filters in the case of at least two in-loop filters, for example after DBF and before SAO. The module 190 is arranged to take as input the same picture as the refinement module 182, i.e., the reconstructed picture if the module 182 is before the in-loop filters, or the partially filtered reconstructed picture if the module 182 is between the in-loop filters.
[0121] FIG. 20 shows a flow diagram of a method for encoding a picture portion into a bitstream according to a particular and non-limiting embodiment. The method begins at step S100. In step S110, a transmitter 1000, such as encoder 100, 101, or 102, accesses a picture portion. As in FIGS. 18 and 19, the accessed picture portion may optionally be mapped before being encoded. In step S120, the transmitter encodes and reconstructs the accessed picture portion to obtain a reconstructed picture portion. To this end, the picture portion may be divided into blocks. Encoding the picture portion includes encoding a block of the picture portion. Encoding the block typically, but not necessarily, includes subtracting a predictor from the block to obtain a block of residuals, transforming the block of residuals into a block of transform coefficients, quantizing the block of coefficients using a quantization step size to obtain a block of quantized transform coefficients, and entropy coding the block of quantized transform coefficients into a bitstream. Reconstructing a block at the encoder side usually, but not necessarily, involves dequantizing and inverse transforming a block of quantized transform coefficients to obtain a block of residuals, and adding a predictor to the block of residuals to obtain a decoded block. The reconstructed picture portion can be filtered by an in-loop filter, e.g., a deblocking / SAO / ALF filter, as in Figures 17-19, or it can be inverse mapped as in Figure 18.
[0122] In step S130, the refinement data is determined, for example by module 190, such that the coding cost of the data (i.e. the coding cost of the refinement data and the refined picture portion) and the rate-distortion cost, calculated as a weighted sum of distortions between the original version of the picture portion, i.e. the accessed image portion possibly mapped as in Figure 19, and the reconstructed picture portion possibly filtered as in Figures 17 and 19 or de-mapped and refined as in Figure 18, are reduced or minimized. More precisely, the refinement data is: - when no mapping is applied before coding as in Figures 17 and 19 without mapping, a rate-distortion cost calculated as a weighted sum of the coding cost of the data (i.e. the coding cost of the improved data and the improved picture portion) and the distortion between an original version of said picture portion and said reconstructed picture portion after improvement with said improved data is determined to be reduced compared to a rate-distortion cost calculated as a weighted sum of the coding cost of the data (i.e. the coding cost of the picture portion without improvement) and the distortion between said original version of said picture portion and said reconstructed picture portion without improvement, - if mapping is applied before coding as in Figure 18 (with mapping, refinement outside the loop) and said refinement is outside the decoding loop, a rate-distortion cost calculated as a weighted sum of coding costs of data (coding costs of refined data and refined picture portion) and distortions between an original version of said picture portion and said reconstructed picture portion after refinement by inverse mapping and refined data is determined so as to be reduced compared to a rate-distortion cost calculated as a weighted sum of coding costs of data and distortions between said original version of said picture portion and said reconstructed picture portion after inverse mapping without refinement, 19 with mapping is applied before coding and said refinement is inside the decoding loop, a rate-distortion cost calculated as a weighted sum of coding cost of data and distortion between a mapped original version of said picture portion and said reconstructed picture portion after refinement with said refinement data is determined so as to be reduced compared to a rate-distortion cost calculated as a weighted sum of coding cost of data and distortion between said mapped original version of said picture portion and said reconstructed picture portion without refinement. The reconstructed picture portion used to determine the refinement data may be an in-loop filtered version or an in-loop partially filtered version of the reconstructed picture portion.
[0123] In a specific and non-limiting embodiment, the refinement data, denoted R, is advantageously modeled by a piecewise linear model (PWL) defined by N pairs (R#idx[k], R#val[k]), k=0 to N-1. Each pair defines a pivot point of the PWL model. Step S130 is detailed in FIGS. 21-26. R#idx[k] is typically a value within the range of the signal under consideration, e.g., in the range 0 to 1023 for a 10-bit signal. Advantageously, R#idx[k] is greater than R#idx[k-1].
[0124] In step S140, the refinement data is coded into the bitstream or signal. A non-limiting example of an embodiment of the syntax is given below. This example considers refining three components. In a variant, the syntax can be coded and applied to only a portion of the components (e.g., only two chroma components). The refinement data is coded in the form of an refinement table.
[0125] [Table 6]
[0126] If refinement#table#new#flag is equal to 0 or refinement#table#flag#luma is equal to 0, no refinement is applied to luma.
[0127] Otherwise, for all pt from 0 to (refinement#table#luma#size-1), calculate the luma refinement value (R#idx[pt], R#val[pt]) as follows: Set R#idx[pt] equal to (default#idx[pt]+refinement#luma#idx[pt]) Set R#val[pt] equal to (NeutralVal+refinement#luma#value[pt])
[0128] A similar process is applied to the cb or cr components.
[0129] For example, NeutralVal=128, which can be used to initialize the value of R#val in step S130, as detailed in Figures 21 and 24. NeutralVal and default#idx[pt] can be default values known on both the encoder and decoder sides, in which case there is no need to transmit these values. In a variant, NeutralVal and default#idx[pt] can be values coded into the bitstream or signal.
[0130] Preferably, default#idx[pt] of pt from 0 to (refinement#table#luma#size-1) is default#idx[pt]=(MaxVal / (refinement#table#luma#size-1))*pt or default#idx[pt]=((MaxVal+1) / (refinement#table#luma#size-1))*pt which corresponds to equidistant indices from 0 to MaxVal or (Max+1), where MaxVal is the maximum value of the signal (eg, 1023 if the signal is represented in 10 bits).
[0131] In a variant, if a mapping is applied and based on the PWL mapping table defined by the pair (map#idx[k], map#val[k]), R#idx[k] for k=0 to N-1 is initialized by map#idx[k]. In other words, default#idx[k] is equal to map#idx[k]. Similarly, in another variant, R#val[k] for k=0 to N-1 can be initialized by map#val[k] for k=0 to N-1.
[0132] The syntax elements refinement#table#luma#size, refinement#table#cb#size, refinement#table#cr#size can be defined by default, in which case there is no need to code them in the stream.
[0133] For equidistant points of the luma PWL model, there is no need to code refinement#luma#idx[pt]. Set refinement#luma#idx[pt] to 0 so that R#idx[pt]=default#idx[pt] for pt between 0 and (refinement#table#luma#size-1). The same applies to the cb or cr tables.
[0134] In a variant, a syntax element can be added for each table to indicate which refinement mode to use (between intra-component (mode 1) and / or inter-component (mode 2)). The syntax element can be signaled at the SPS, PPS, slice, tile, or CTU level.
[0135] In a variant, a syntax element can be added for each table to indicate whether the table is applied as a multiplicative or additive operator. The syntax element can be signaled at the level of SPS, PPS, slice, tile, or CTU.
[0136] In one embodiment, the refinement table is not coded into the bitstream or signal, but instead a default inverse mapping table (corresponding to the inverse of the mapping table used by mapping 105 in Figures 18 and 19) is modified with the refinement table, and the modified inverse mapping table is coded into the bitstream or signal.
[0137] In one embodiment, refinement tables are coded only for pictures at a low temporal level, for example, only for pictures at temporal level 0 (the lowest level in the temporal coding hierarchy).
[0138] In one embodiment, refinement tables are coded only for random access pictures, such as intra pictures.
[0139] In one embodiment, refinement tables are coded only for high quality pictures that correspond to an average QP over the pictures below a given value.
[0140] In one embodiment, the refinement table is coded only if the coding cost of the refinement table relative to the coding cost of the full picture is below a given value.
[0141] In one embodiment, the refinement table is coded only if the rate-distortion gain compared to not coding the refinement table exceeds a given value. For example, the following rules may apply: · Code the table if the rate-distortion cost gain is greater than (0.01*width*height), where width and height are the dimensions of the component under consideration. · Code the table if the rate-distortion cost gain exceeds (0.0025*initRD), where initRD is the rate-distortion cost without applying component refinements. When the rate-distortion cost is based on squared error distortion, this roughly corresponds to a minimum PSNR gain of 0.01 dB. If the given value is set to (0.01*initRD), this roughly corresponds to a minimum PSNR gain of 0.05 dB when the rate-distortion cost is based on squared error distortion. The values listed are examples only and can be modified.
[0142] Returning to Figure 20, in optional step S150, refinement data is applied to the reconstructed picture portion that is possibly filtered as in Figure 19. To do this, a look-up table LutR is determined from pairs of points (R#idx[pt], R#val[pt]) of the PWL for pt=0 to N-1. For example, LutR is determined by linear interpolation between each pair of points (R#idx[pt], R#val[pt]) and (R#idx[pt+1], R#val[pt+1]) of the PWL as follows: pt=0 to N-2 From idx=R#idx[pt] to (R#idx[pt+1]-1), LutR[idx]=R#val[pt]+(R#val[pt+1]-R#val[pt])*(idx-R#idx[pt]) / (R#idx[pt+1]-R#idx[pt])
[0143] In a modified form, LutR is determined as follows: pt=0 to N-2 From idx=R#idx[pt] to (R#idx[pt+1]-1), LutR[idx]=(R#val[pt]+R#val[pt+1]) / 2 Two examples of improved modes are: Mode 1 - Intra-component refinement. In mode 1, the refinement is performed independently of the other components. The signal Srec(p) is refined as follows: Sout(p)=LutR[Srec(p)] / NeutralVal*Srec(p) where Srec(p) is the reconstructed signal to be improved at position p in the picture portion, corresponding to the signal resulting from the in-loop filter or possibly from the inverse mapping, and Sout(p) is the improved signal. Here, it is considered that the signals Srec and Sout use the same bit depth. If different bit depths are used for both signals (i.e. Bout for Sout and Brec for Srec), a scaling factor related to the difference in bit depth can be applied. For example, if the bit depth Bout of Sout is higher than the bit depth Brec of Srec, the following formula is applied: Sout(p)=2 (Bout-Brec) *LutR[Srec(p)] / NeutralVal*Srec(p) The scaling factor can be directly integrated into the LutR value, for example if the bit depth Bout of Sout is lower than the bit depth Brec of Srec, then the formula applies as follows: Sout(p)=LutR[Srec(p)] / NeutralVal*Srec(p) / 2 (Brec-Bout) The scaling factor can be directly integrated into the value of LutR. Advantageously, mode 1 is used for the luma component. Mode 2 - Inter-component refinement. In mode 2 the refinement is performed on one component C0 and depends on another component C1. The signal Srec#C0(p) is refined as follows: Sout(p)=offset+LutR[Srec#C1(p)] / NeutralVal*(Srec#C0(p)-offset) where offset is set to, for example, (MaxVal / 2), Srec#C0(p) is the reconstructed signal of the component C0 to be improved, and p is the sample position within the picture part. MaxVal is the maximum value of the signal Srec#C0, and (2 B-1), where B is the bit depth of the signal. Srec#C1(p) is the reconstructed signal of component C1. Srec#C1(p) can be the signal resulting from an in-loop filter or from an inverse mapping. After filtering by the in-loop filter, Srec#C1(p) can also be further filtered, for example with a low-pass filter. Here we consider that signals Srec#C0, Srec#C1, and Sout use the same bit depth. Advantageously, mode 2 can be applied to the chroma components and depends on the luma component. If the bit depth Bout of Sout is higher than the bit depth Brec of Srec#C0 and Srec#C1, then the formula is adapted as follows: Sout(p)=2 (Bout-Brec) *(offset+LutR[Srec#C1(p)] / NeutralVal*(Srec#C0(p)-offset)) If the bit depth Bout of Sout is lower than the bit depth Brec of Srec#C0 and Srec#C1, the formula applies as follows: Sout(p)=(offset+LutR[Srec#C1(p)] / NeutralVal*(Srec#C0(p)-offset)) / 2 (Bout-Brec) Rounding and clipping between the minimum and maximum signal values (typically between 0 and 1023 for a 10-bit signal) is finally applied to the refined value Sout(p).
[0144] In the above embodiment, the refinement is applied as a multiplicative operator. In a variant, the refinement is applied as an additive operator. In this case, in mode 1, the signal Srec(p) is refined as follows: 〇 Sout(p)=LutR[Srec(p)] / NeutralVal+Srec(p) In this case, in mode 2, the signal Srec#C0(p) is improved as follows: 〇 Sout(p)=LutR[Srec#C1(p)] / NeutralVal+Srec#C0(p)
[0145] For the chroma components Cb and Cr, an example of values for an inter-component refinement table for NeutralVal=64 and N=17 is shown below:
[0146] [Table 7]
[0147] [Table 8]
[0148] Returning to FIG. 20, the method ends in step S160.
[0149] Figure 21 shows a flow diagram illustrating an example of further details regarding step S130. The refinement data can be determined on a given picture region A, e.g., a slice, tile, or CTU, or on the whole picture. The refinement data R can advantageously be modeled by a piecewise linear model (PWL) defined by N pairs (R#idx[k], R#val[k]), k=0 to N-1. An example of a PWL model with N=6 (k∈[0,1,2,3,5]) is shown in Figure 22. Each pair defines a pivot point of the PWL model.
[0150] The values of R#idx and R#val are initialized (step S1300). Typically, R#idx[k] is initialized for k=0 to N-1, so that there are equidistant intervals between consecutive indices, i.e., (R#idx[k+1]-R#idx[k])=D, where D=Range / (N-1), and Range is the range of the signal to be improved (e.g., 1024 for a signal represented by 10 bits). In one example, N=17 or 33. The value of R#val is initialized with a value of NeutralVal, for example, 128, which is determined so that the improvement does not change the signal. In a modified form, when mapping is applied based on the PWL mapping table defined by the pair (map#idx[k], map#val[k]), (R#idx[k], R#val[k]) for k=0 to N-1 is initialized by (map#idx[k], map#val[k]).
[0151] In the embodiment of Figure 21, only the value of R#val[k] is determined. The value of R#idx[k] is fixed to its initial value. Using R initialized in S1300, the initial rate-distortion cost initRD is calculated (step 1301). The initial rate-distortion cost initRD is calculated as follows: initRD=L*Cost(R)+Σ p in A dist(Sin(p), Sout(p)) (Equation 1) however - R is the improved data initialized by S1300, A is the picture area to be refined, - Sin(p) is as follows, i.e. If no mapping is applied, Sin(p) is the sample value of pixel p in the original picture domain, If the mapping is applied and the refinement is outside the loop (Figure 18), Sin(p) is the sample value of pixel p in the original picture domain, If mapping is applied and refinement is in a loop (Figure 19), Sin(p) is the sample value of pixel p in the mapped original picture region. It is so determined, Sout(p) is the sample value of pixel p in the refined picture area, - dist(x,y) is the distortion between sample value x and sample value y, for example, the distortion is the squared error (xy) 2 and other possible distortion functions are the absolute difference |xy|, or distortions based on subjective metrics such as SSIM (Z. Wang, A.C. Bovik, H.R. Sheikh and E.P. Simoncelli, “Image quality assessment: From error visibility to structural similarity,” IEEE Transactions on Image Processing, vol. 13, no. 4, pp. 600-612, April 2004), or modifications of SSIM can also be used. Cost(R) is the coding cost for coding the refinement data R and the refined picture region, L is a value related to the picture area A. L is advantageously 2 (QP / 6) , where QP is the quantization parameter if a single quantization parameter value is used in picture region A, or if different quantization parameter values are used in picture region A, QP represents the quantization parameter applied to picture region A. For example, QP is the average of the QPs used in the blocks of region A. If the value of R_val is initialized with the value of NeutralVal, the initial rate-distortion cost initRD can be calculated using the coding cost of the picture region without the coding cost of the refinement data.
[0152] In step S1302, the parameter bestRD is initialized to initRD. Next, in step S1303, refinement data R is determined. In step S1304, a loop is performed over the indices pt of successive pivot points R of the PWL model. In step S1305, the parameters bestVal and initVal are initialized to R#val[pt]. In step S1306, a loop is performed over various values of R#val[pt], from (initVal-Val0) to (initVal+Val1), where Val0 and Val1 are default parameters. Typical values are Val0=Val1=NeutralVal / 4. In step S1307, the rate-distortion cost curRD is calculated using equation 1 with the current R (along with the current R#val[pt]). In step S1308, curRD is compared with bestRD. If curRD is lower than bestRD, bestRD is set to curRD and bestValue is set to R#val[pt]. Otherwise, the method continues at step S1310. At step S1310, it is determined whether the loop over the value of R#val[pt] has finished. If the loop has finished, then at step S1311 R#val[pt] is set to bestValue. At step S1312, it is determined whether the loop over the value of pt has finished. If the loop has finished, then the refinement data to output is the current R.
[0153] Step S1303 can be repeated n times, where n is a fixed integer, for example, n = 3. Figure 23 shows the determination of R#val. The dashed line represents the PWL after updating R#val.
[0154] Figure 24 shows a modification of the process of Figure 21. In the embodiment of Figure 24, only the value of R#idx[k] is determined. The value of R#val[k] is fixed at its initial value. The process uses the initial R#idx and R#val data (e.g., resulting from the method of Figure 21) as input. In step S1401, the initial rate-distortion cost initRD is calculated using (Equation 1). Then, in step S1402, the parameter bestRD is initialized to initRD. In step S1403, the refinement data R is determined. In step S1404, a loop is performed over the indices pt of successive pivot points R of the PWL model. In step S1405, the parameters bestIdx and initIdx are initialized to R#idx[pt]. In step S1406, a loop is performed over various values of R#idx[pt] (from initIdx-idxVal0 to initIdx+idxVal1), where idxVal0 and idxVal1 are specified values, e.g., idxVal0=idxVal1=D / 4, where D=Range / (N-1), and Range is the range of the signal to be improved (e.g., 1024 for a signal represented by 10 bits). In step S1407, the rate-distortion cost curRD is calculated using equation 1 together with the current R (along with the current R#idx[pt]). In step S1408, curRD is compared with bestRD. If curRD is lower than bestRD, bestRD is set to curRD and bestIdx is set to R#idx[pt]. Otherwise, the method continues to step S1310. In step S1410, it is determined whether the loop over the values of R#idx[pt] has finished. If the loop has finished, in step S1411 R#idx[pt] is set to bestIdx. In step S1412, it is checked whether the loop on the value of pt has finished. If the loop has finished, the refined data to be output is the current R. Step S1403 can be repeated n times, where n is a fixed integer, for example, n=3.
[0155] Figure 25 shows the determination of R#idx. The dashed line represents the PWL after updating R#idx.
[0156] Figure 26 shows the determination of both R#idx and R#val. The dashed line represents the PWL after updating R#val and R#idx.
[0157] 27 illustrates an example of the architecture of a receiver 2000 configured to decode pictures from a bitstream or signal to obtain decoded pictures, according to a particular and non-limiting embodiment. The receiver 2000 includes one or more processors 2005, which may include, for example, a CPU, a GPU, and / or a DSP (an English acronym for digital signal processor), along with built-in memory 2030 (e.g., RAM, ROM, and / or EPROM). The receiver 2000 includes one or more communication interfaces 2010 (e.g., keyboard, mouse, touchpad, webcam), each adapted to display output information and / or allow a user to input commands and / or data (e.g., decoded pictures), and a power source 2020, which may be external to the receiver 2000. The receiver 2000 may also include one or more network interfaces (not shown). A decoder module 2040 represents a module that may be included within a device to perform decoding functions. Additionally, the decoder module 2040 may be implemented as a separate element of the receiver 2000 or may be incorporated within the processor 2005 as a combination of hardware and software, as known to those skilled in the art.
[0158] The bitstream or signal can be obtained from a source, according to different embodiments, including but not limited to: - Local memory, e.g. video memory, RAM, flash memory, hard disk, - a storage interface, e.g. an interface to a mass storage, ROM, optical disk, or magnetic support; a communication interface, such as a wired interface (e.g. a bus interface, a wide area network interface, a local area network interface) or a wireless interface (such as an IEEE 802.11 interface or a Bluetooth interface), and - Image capture circuitry (e.g. sensors such as CCD (i.e. charge coupled device) or CMOS (i.e. complementary metal oxide semiconductor)) It could be.
[0159] According to different embodiments, the decoded pictures can be transmitted to a destination, for example a display device. By way of example, the decoded pictures are stored in a remote or local memory, for example a video memory, or a RAM, or a hard disk. In variants, the decoded pictures are transmitted to a storage interface, for example an interface to a mass storage, a ROM, a flash memory, an optical disk, or a magnetic support, and / or transmitted over a communications interface, for example an interface to a point-to-point link, a communications bus, a point-to-multipoint link, or a broadcast network.
[0160] According to a specific and non-limiting embodiment, receiver 2000 further includes a computer program stored in memory 2030. The computer program includes instructions that, when executed by receiver 2000, specifically by processor 2005, enable the receiver to perform the decoding method described with respect to FIG. 31. According to a variant, the computer program is stored external to receiver 2000 on a non-transitory digital data support, for example, on an external storage medium such as a HDD, CD-ROM, DVD, read-only and / or DVD drive, and / or DVD read / write drive, all known in the art. Receiver 2000 therefore includes a mechanism for reading the computer program. Furthermore, receiver 2000 can access one or more Universal Serial Bus (USB) type storage devices (e.g., "memory sticks") via corresponding USB ports (not shown).
[0161] According to one or more non-limiting example embodiments, the receiver 2000 may include, but is not limited to: - mobile devices, - communication devices, - gaming consoles, - set-top boxes, - TV sets, - a tablet (or tablet computer), - laptop, - Video players, such as Blu-ray players, DVD players, - Display, and - Decryption chip or decryption device / equipment It could be.
[0162] Figure 28 shows a block diagram of an example embodiment of a video decoder 200, e.g., of the HEVC type, adapted to perform the decoding method of Figure 31. The video decoder 200 is an example of a receiver 2000 or part of such a receiver 2000. In the illustrated example embodiment of the decoder 200, a bitstream or signal is decoded by elements of the decoder as described below. The video decoder 200 generally performs a decoding pass that is the inverse of the encoding pass described in Figure 17, which performs video decoding as part of encoding the video data.
[0163] Specifically, the decoder input includes a video bitstream or signal, such as may be generated by video encoder 100. The bitstream or signal is first entropy decoded (230) to obtain transform coefficients, motion vectors, and other coded information, e.g., refinement data. The transform coefficients are dequantized (240) and inverse transformed (250) to decode the residual. The decoded residual is combined (255) with a predicted block (also known as a predictor) to obtain a decoded / reconstructed picture block. The predicted block may result from intra-prediction (260) or motion-compensated prediction (i.e., inter-prediction) (275) (270). As noted above, AMVP and merge mode techniques can be used during motion compensation, which may use interpolation filters to calculate interpolated values for sub-integer samples of the reference block. An in-loop filter (265) is applied to the reconstructed picture. The in-loop filter may include a deblocking filter and an SAO filter. The filtered picture is stored in a reference picture buffer (280). The possibly filtered reconstructed picture is refined 290. The refinement is applied outside the decoding loop as a post-process.
[0164] Figure 29 shows a modification 201 of the video decoder 200 of Figure 28. Modules in Figure 29 that are identical to modules in Figure 28 are labeled with the same reference numbers and will not be described further. The filtered reconstructed picture, i.e., the output of the in-loop filter, is inversely mapped (285). The inverse mapping (285) is the reverse process of the mapping (105) applied at the encoder side. The inverse mapping can use an inverse mapping table decoded from the bitstream or signal, or a pre-defined inverse mapping table. The inverse mapped picture is refined (290) using refinement data decoded (230) from the bitstream or signal.
[0165] A variant merges the inverse mapping and refinement into a single module that applies the inverse mapping using an inverse mapping table decoded from the bitstream or signal, and the inverse mapping table is modified in the encoder to take the refinement data into account. A variant applies a lookup table LutComb to perform the inverse mapping and refinement process for a given component being processed, this lookup table being constructed as the concatenation of a lookup table LutInvMap derived from the mapping table and a lookup table derived from the refinement table LutR: LutComb[x]=LutR[LutInvMap[x]], x=0 to MaxVal In this embodiment, the refinement process is outside the decoding loop, and is therefore applied in the decoder only as a post-process.
[0166] Figure 30 shows a variation 202 of the video decoder 200 of Figure 28. Modules in Figure 30 that are identical to modules in Figure 28 are labeled with the same reference numbers and will not be described further. Refinement data is decoded from the bitstream or signal (230). The decoded refinement data is used to refine the filtered reconstructed picture (290). The refined picture, rather than the filtered reconstructed picture, is stored in the reference picture buffer (280). The refinement module 290 can be inserted in various locations. The refinement module 290 can be inserted before the in-loop filter or between the in-loop filters in the case of at least two in-loop filters, for example, after DBF and before SAO. The refined picture can optionally be inverse mapped (285). In this embodiment, the refinement process is within the decoding loop.
[0167] FIG. 31 shows a flow diagram of a method for decoding a picture from a bitstream or signal according to a particular and non-limiting embodiment. The method begins at step S200. In step S210, a receiver 2000, such as decoder 200, accesses the bitstream or signal. In step S220, the receiver decodes a picture portion from the bitstream or signal to obtain a decoded picture portion by decoding a block of the picture portion. Decoding a block typically, but not necessarily, involves entropy decoding a portion of the bitstream or signal representing the block to obtain a block of transform coefficients, dequantizing and inverse transforming the block of transform coefficients to obtain a block of residuals, and adding a predictor to the block of residuals to obtain a decoded block. The decoded picture portion may then be filtered by an in-loop filter, as in FIGS. 28-30, and possibly reverse-mapped, as in FIG. 29. In step S230, refinement data is decoded from the bitstream or signal. This step is the inverse of encoding step S140. All the variations and embodiments described for step S140 also apply to step S230. In step S240, the decoded picture is refined. This step is identical to the encoder-side refinement step S150.
[0168] Another example embodiment is shown in Figure 32, which is a method for encoding video data that includes providing improved mode functionality. In Figure 32, video data, such as data contained in a digital video signal or bitstream, may be processed at 3210 to identify costs associated with encoding the video data, such as a picture portion, using a block-by-block improved mode described herein and using a mode other than the improved mode, e.g., no improvement. The video data is then encoded at 3220 based on the costs. For example, a cost, such as a rate-distortion cost, for both encoding using the improved mode and encoding using the no-improvement mode may be obtained. If using the improved mode results in a rate-distortion cost that indicates an improvement in the coding, then the improved mode may be used for encoding. A rate-distortion cost that indicates no improvement, with limited improvement that does not justify the added complexity of the improved mode, may result in not using the improved mode or using a different mode. Processing the video data to obtain the costs and to encode the video data may be performed, for example, as described above with respect to one or more embodiments described herein. The encoded video data is output from 3220.
[0169] 33 illustrates an example embodiment of a method for decoding video data. Encoded video data, e.g., a digital data signal or bitstream, is processed at 3310 to obtain an indication of the use of an refinement mode, e.g., a block-by-block refinement mode, during the encoding of the video data according to one or more embodiments described herein. Then, at 3320, the encoded video data, e.g., the encoded picture portion, is decoded based on the indication (e.g., decoding as described herein based on the refinement mode used during encoding). Processing the video data to obtain the cost and to decode the video data may be performed, e.g., as described above with respect to one or more embodiments described herein. Decoded video data may be output from 3320.
[0170] Throughout this disclosure, various implementations include decoding. As used herein, "decoding" may encompass all or part of the processes performed on a received encoded sequence to produce a final output suitable for display, for example. In various embodiments, such processes include one or more of the processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also or instead include the processes performed by decoders of various implementations described herein, such as extracting a picture from tiled (packed) pictures, determining an upsampling filter to use and upsampling the picture, and flipping the picture back to its intended orientation.
[0171] As a further example, in some embodiments, "decoding" refers to entropy decoding only, in other embodiments "decoding" refers to differential decoding only, and in other embodiments "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to refer specifically to a subset of operations or to a broader decoding process generally will be clear based on the context of the specific description and is believed to be well understood by one of ordinary skill in the art.
[0172] Additionally, various implementations include encoding. Similar to the discussion above regarding "decoding," "encoding," as used herein, may encompass all or part of the processes performed on an input video sequence to, for example, produce an encoded bitstream or signal. In various embodiments, such processes include one or more of the processes typically performed by an encoder, such as segmentation, differential encoding, transforming, quantizing, and entropy encoding. In various embodiments, such processes also or instead include processes performed by the encoders of various implementations described herein.
[0173] As a further example, in some embodiments, "encoding" refers only to entropy encoding, in other embodiments "encoding" refers only to differential encoding, and in other embodiments "encoding" refers to a combination of differential and entropy encoding. Whether the phrase "encoding process" is intended to refer specifically to a subset of operations or to a broader encoding process generally will be clear based on the context of the specific description and is believed to be well understood by one of ordinary skill in the art.
[0174] It should be noted that the syntax elements used herein are descriptive terms, so they do not preclude the use of other syntax element names.
[0175] Where a drawing is shown as a flow diagram, it should be understood that the drawing also provides a block diagram of the corresponding apparatus. Similarly, where a drawing is shown as a block diagram, it should be understood that the drawing also provides a flow diagram of the corresponding method / process.
[0176] Various embodiments refer to rate-distortion optimization. Specifically, during the encoding process, a balance or trade-off between rate and distortion is typically considered, often given a computational complexity constraint. Rate-distortion optimization is typically formulated as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. There are various approaches to solving the rate-distortion optimization problem. For example, these approaches may involve a thorough evaluation of the coding cost and the associated distortion of the reconstructed signal after coding and decoding, and may be based on a comprehensive testing of all coding options, including all modes or coding parameter values considered. Faster approaches can also be used to reduce coding complexity, particularly by calculating approximate distortion based on a predicted or prediction residual signal rather than the reconstructed signal. A hybrid of these two approaches can also be used, such as by using approximate distortion for only some coding options and full distortion for other coding options. Other approaches evaluate only a subset of coding options. More generally, many approaches use any of a variety of techniques for performing optimization, but the optimization does not necessarily involve a thorough evaluation of both the coding cost and the associated distortion.
[0177] Implementations and aspects described herein may be implemented by, for example, a method or process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single type of implementation (e.g., only discussed as a method), the implementation of the discussed features may also be implemented in other forms (e.g., an apparatus or a program). An apparatus may be implemented by, for example, appropriate hardware, software, and firmware. A method may be implemented by, for example, a processor, where a processor refers generally to processing devices, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end users.
[0178] Reference to "one embodiment," or "an embodiment," or "one implementation," or "an implementation," as well as other variants thereof, means that a particular feature, structure, characteristic, etc. described in connection with an embodiment is included in at least one embodiment. Thus, appearances of the phrases "in one embodiment," or "in an embodiment," or "in one implementation," or "in an implementation," as well as any other variants, appearing in various places throughout this specification do not necessarily all refer to the same embodiment.
[0179] Additionally, this specification may refer to "obtaining" various pieces of information. Obtaining information may include, for example, one or more of determining information, estimating information, calculating information, predicting information, or retrieving information from memory.
[0180] Additionally, this specification may refer to "accessing" various pieces of information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, replicating information, calculating information, determining information, predicting information, or estimating information.
[0181] Additionally, this specification may refer to "receiving" various pieces of information. Receiving is intended to be broad, similar to "accessing." Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Furthermore, "receiving" typically involves some form of operation, such as, for example, storing information, processing information, transmitting information, moving information, duplicating information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0182] For example, the use of " / ", "and / or", and "at least one of" in the cases of "A / B", "A and / or B", and "at least one of A and B" is intended to encompass selecting only the first listed (A) option, or selecting only the second listed (B) option, or selecting both (A and B) options. As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such phrases are intended to encompass selecting only the first listed (A) option, or selecting only the second listed (B) option, or selecting only the third listed (C) option, or selecting only the first and second listed options (A and B), or selecting only the first and third listed options (A and C), or selecting only the second and third listed options (B and C), or selecting all three options (A, B, and C). As will be apparent to those skilled in the art, this representation can be extended to include as many items as are listed.
[0183] Furthermore, as used herein, the term "signaling" refers to, among other things, indicating something to a corresponding decoder. For example, in certain embodiments, an encoder signals to improve a particular one of a plurality of parameters. In this manner, the same parameters are used at both the encoder and decoder sides in one embodiment. Thus, for example, an encoder can transmit a particular parameter to a decoder (explicit signaling), allowing the decoder to use the same particular parameter. Conversely, if the decoder already has the particular parameter along with other parameters, signaling can be used without transmission simply to allow the decoder to know and select the particular parameter (implicit signaling). By avoiding transmitting any actual functionality, bit savings are realized in various embodiments. It should be understood that signaling can be implemented in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder. While the above content relates to the verb form of the word "signal," the word "signal" can also be used as a noun herein.
[0184] As will be apparent to one skilled in the art, implementations can produce a variety of signals formatted to carry information that can be, for example, stored or transmitted. Information can include, for example, instructions for performing a method or data produced by one of the described implementations. For example, a signal can be formatted to carry a bit stream or signal of the described embodiments. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.
[0185] Various embodiments have been described. The embodiments may include any of the following features or entities, alone or in any combination, across various different claim categories and types: Applying refinements that are independent of neighboring reconstructed samples. Applying a refinement based on a global function having one or more refinement parameters transmitted in the bitstream. · Providing modes within the encoder and / or decoder that include inter-component refinement of chroma components. Providing modes in the encoder and / or decoder that include intra-component refinement of the luma component. · To be able to activate block-wise refinement within the decoder and / or encoder. · Applying refinements by using several functions for specific components and selecting the refinement function block by block. Allowing the decoder and / or encoder to select, on a block-by-block basis, refinement parameters to apply to the block from a set of possible parameters coded in the bitstream or signal. Applying block-wise refinements to the reconstructed signal. Include a refinement step as an in-loop or out-of-loop filter to improve the reconstructed signal after decoding. Within the decoder and / or encoder, basing the refinement on refinement tables coded into the bitstream or signal. · Providing a reduction in table coding costs by avoiding redundant neutral values. Inserting syntax elements into the signaling that enable the decoder to make the improvements described herein. · Include identifiers related to refinement tables, such as table pair identifiers, within syntax elements. · Limiting the size of refinement tables to reduce coding costs. · Including one or more syntax elements within the syntax element that indicate whether refinement information for the current block is replicated from neighboring blocks. Based on these syntax elements, select the refinements to apply in the decoder. A bitstream or signal containing one or more of the syntax elements described or variations thereof. Inserting syntax elements into the signaling that allow the decoder to refine it in a way that corresponds to the way used by the encoder. Creating and / or transmitting, and / or receiving and / or decoding bitstreams or signals that include one or more of the listed syntax elements or variations thereof. A TV, set-top box, mobile phone, tablet, or other electronic device incorporating an improvement according to any of the described embodiments. A TV, set-top box, mobile phone, tablet, or other electronic device that implements an improvement according to any of the described embodiments and displays the resulting image (e.g., using a monitor, screen, or other type of display). A TV, set-top box, mobile phone, tablet, or other electronic device that tunes to a channel (e.g., using a tuner) to receive a signal containing encoded images and that is improved according to any of the described embodiments. A TV, set-top box, mobile phone, tablet, or other electronic device that receives a signal containing encoded images wirelessly (e.g., using an antenna) and that is improved according to any of the described embodiments. A computer program product storing program code which, when executed by a computer, implements an improvement according to any of the described embodiments. A non-transitory computer-readable medium containing executable program instructions that cause a computer executing the instructions to implement an improvement according to any of the described embodiments.
[0186] Various other generalized and specialized embodiments are also supported and contemplated throughout this disclosure. For example, at least one example embodiment includes a method of encoding that includes: (i) obtaining a cost for encoding a picture portion based on applying an improvement mode to a block-by-block reconstructed signal, where the improvement mode is based on an improvement parameter; and (ii) obtaining a cost for encoding the picture portion using a mode other than the improvement mode, and encoding the picture portion based on the cost.
[0187] Another example of at least one embodiment includes an apparatus including one or more processors configured to obtain a cost for encoding a picture portion based on applying an improvement mode to a block-by-block reconstructed signal, where the improvement mode is based on an improvement parameter, and a cost for encoding the picture portion without using the improvement mode, and to encode the picture portion based on the cost.
[0188] Another example of at least one embodiment includes a method of decoding that includes obtaining an indication of an improvement mode that has been applied to a block-by-block reconstructed signal during encoding of the picture portion, the improvement mode being based on an improvement parameter, and decoding the encoded picture portion based on the indication.
[0189] Another example of at least one embodiment includes an apparatus including one or more processors configured to obtain an indication of an improvement mode that has been applied to a block-by-block reconstructed signal during encoding of the encoded picture portion, the improvement mode being based on an improvement parameter, and decoding the encoded picture portion based on the indication.
[0190] Another example of at least one embodiment includes a signal formatted to include data representing an encoded picture portion and data providing an indication of an enhancement mode that has been applied to a block-by-block reconstructed signal based on enhancement parameters during encoding of the encoded picture portion.
[0191] Another example of at least one embodiment includes a bitstream formatted to include data representing an encoded picture portion and data providing an indication of an enhancement mode that has been applied to a block-by-block reconstructed signal based on enhancement parameters during encoding of the encoded picture portion.
[0192] In at least one embodiment described herein that includes an improved mode, the improved mode may include ingredient improvement.
[0193] In at least one embodiment described herein that includes an improved mode, the improved mode can include at least one of an inter-component improvement or an intra-component improvement.
[0194] In at least one embodiment described herein that includes inter-component or intra-component improvements, the inter-component improvements can include inter-component chroma improvements, and the intra-component improvements can include intra-component luma improvements.
[0195] In at least one embodiment described herein that includes a refinement mode, the refinement mode may include allowing block-by-block selection of refinement parameters.
[0196] In at least one embodiment described herein that includes refinement parameters, the refinement parameters may include one or more refinement parameters contained in a refinement table.
[0197] In at least one embodiment described herein that involves encoding a picture portion based on applying an improvement mode and improvement parameters, an improvement parameter may be selected for each block from a plurality of improvement parameters based on the cost improvement when encoding the picture portion using the improvement mode.
[0198] According to another embodiment, a method for encoding video data is presented that includes obtaining (i) a cost for encoding a picture portion using a block-by-block refinement mode, where the refinement mode is based on refinement parameters, and (ii) a cost for encoding the picture portion using a mode other than the refinement mode, and encoding the picture portion based on the cost.
[0199] According to another embodiment, a method for decoding video data is presented, comprising obtaining an indication of the use of a block-by-block refinement mode based on refinement parameters during encoding of an encoded picture portion, and decoding the encoded picture portion based on the indication.
[0200] According to another embodiment, an apparatus for decoding video data is presented, the apparatus including one or more processors configured to obtain an indication of use of a block-by-block refinement mode based on refinement parameters during encoding of an encoded picture portion, and to decode the encoded picture portion based on the indication.
[0201] According to another embodiment, a signal format is presented that includes data representing an encoded picture portion and data providing an indication of the use of a block-by-block refinement mode based on refinement parameters during encoding of the encoded picture portion.
[0202] According to another embodiment, a bitstream is presented that is formatted to include encoded video data, the encoded video data being encoded by obtaining (i) a cost for encoding a picture portion using a block-by-block refinement mode, the refinement mode being based on refinement parameters, and (ii) a cost for encoding the picture portion using a mode other than the refinement mode, and encoding the picture portion based on the cost.
[0203] According to another embodiment, the modification modes described herein may include modification of ingredients.
[0204] According to another embodiment, the improvement modes described herein can include component improvement, where component improvement includes at least one of cross-component improvement, intra-component improvement, and intra-component improvement.
[0205] According to another embodiment, the improvement modes described herein may include inter-component chroma improvement.
[0206] According to another embodiment, the improvement modes described herein may include intra-component luma improvement.
[0207] According to another embodiment, the refinement modes described herein may include allowing block-by-block selection of refinement parameters.
[0208] According to another embodiment, the refinement modes described herein may be based on refinement parameters, including one or more chroma refinement parameters contained in a chroma refinement table.
[0209] According to another embodiment, a method, apparatus, or signal according to the present disclosure may include encoding and / or decoding video data using an improved mode based on improved parameters, wherein the encoding and / or decoding includes encoding and / or decoding picture portions and improved parameters based on an improved cost when encoding and / or decoding the picture portions using the improved mode.
[0210] According to another embodiment, a method, apparatus, or signal according to the present disclosure may include encoding and / or decoding video data based on using an improved mode during encoding and / or decoding, and determining a cost associated with using the improved mode, where the cost may include a rate-distortion cost, and determining the cost may include determining a rate-distortion cost when encoding the picture portion using the improved mode.
[0211] Generally, at least one embodiment of an apparatus for encoding and / or decoding video data may include the equipment described herein and may include at least one of: (i) an antenna configured to receive a signal, the signal including data representing an image; (ii) a band limiter configured to limit the received signal to a frequency band including the data representing the image; or (iii) a display configured to display the image.
[0212] Generally, at least one embodiment of the device for encoding and / or decoding video data may include at least one of a television, a mobile phone, a tablet, a set-top box, and a gateway device.
[0213] One or more embodiments herein also provide a computer-readable storage medium or computer program product storing instructions for encoding or decoding video data according to the methods or apparatus described herein.Embodiments herein also provide a computer-readable storage medium or computer program product storing a bitstream generated according to the methods or apparatus described herein.Embodiments herein also provide methods and apparatus for transmitting or receiving a bitstream generated according to the methods or apparatus described herein.
Claims
1. decoding instructions at a block level of a picture portion, the instructions enabling a block-by-block selection of an improvement parameter table from a plurality of improvement parameter tables associated with an inter-component improvement mode to be applied to a reconstructed version of a chroma component of the picture portion; using the selected refinement parameter table to obtain one or more chroma refinement parameters for the block; and Refining the picture portion in the inter-component refinement mode, the inter-component refinement mode being applied block-by-block of the picture portion based on the obtained one or more chroma refinement parameters. A method comprising:
2. The method of claim 1 , wherein the one or more chroma refinement parameters are pivot points of a piecewise linear model.
3. The method of claim 2 , further comprising constructing a lookup table using the one or more chroma improvement parameters.
4. The method of claim 3 , wherein the reconstructed version of the chroma components of the picture portion is refined using the constructed look-up table.
5. The method of claim 1 , wherein the plurality of refinement parameter tables comprises a plurality of pairs of chroma refinement tables.
6. The method of claim 1 , further comprising decoding the refinement parameter tables from a slice header of the picture portion.
7. The method of claim 6 , wherein the plurality of refinement parameter tables includes one or more refinement parameters for refining the one or more chroma refinement parameters.
8. 8. The method of claim 7, wherein the one or more refinement parameters include an index for updating an index of the one or more chroma refinement parameters and a refinement value for updating a value of the one or more chroma refinement parameters.
9. The method of claim 1 , wherein a given value of the indication enables the inter-component refinement mode to be enabled for the block.
10. The method of claim 1 , wherein enabling or disabling the inter-component refinement mode for the block is inferred from an SAO parameter of the block.
11. 7. The method of claim 6, further comprising: decoding, for the picture portion, a first flag indicating whether at least one first improvement parameter table is coded for a first chroma component of the picture portion, and a second flag indicating whether at least one second improvement parameter table is coded for a second chroma component of the picture portion, wherein the plurality of improvement parameter tables are further decoded based on at least one of the first flag or the second flag.
12. A non-transitory computer readable medium storing executable program instructions for causing a computer executing the instructions to perform the method of claim 1.
13. decoding instructions at a block level of a picture portion, the instructions enabling a block-by-block selection of an improvement parameter table from a plurality of improvement parameter tables associated with an inter-component improvement mode to be applied to a reconstructed version of a chroma component of the picture portion; using the selected refinement parameter table to obtain one or more chroma refinement parameters for the block; and enhancing the picture portion in the inter-component refinement mode, the inter-component refinement mode being applied block by block of the picture portion based on the obtained one or more chroma refinement parameters. one or more processors configured to perform Including, equipment.
14. - encoding instructions at a block level of a picture portion, the instructions enabling a block-by-block selection of an improvement parameter table from a plurality of improvement parameter tables associated with an inter-component improvement mode to be applied to a reconstructed version of a chroma component of the picture portion; using the selected refinement parameter table to obtain one or more chroma refinement parameters for the block; and Refining the picture portion in the inter-component refinement mode, the inter-component refinement mode being applied block-by-block of the picture portion based on the obtained one or more chroma refinement parameters. A method comprising:
15. 15. The method of claim 14, wherein the refinement parameter table is selected for each block from the plurality of refinement parameter tables based on an improvement in a cost associated with encoding the picture portion using the inter-component refinement mode.
16. 15. The method of claim 14, further comprising: encoding the plurality of refinement parameter tables in a slice header of the picture portion, the plurality of refinement parameter tables including one or more refinement parameters for refining the one or more chroma refinement parameters.
17. A non-transitory computer readable medium storing executable program instructions for causing a computer executing the instructions to perform the method of claim 14.
18. - encoding instructions at a block level of a picture portion, the instructions enabling a block-by-block selection of an improvement parameter table from a plurality of improvement parameter tables associated with an inter-component improvement mode to be applied to a reconstructed version of a chroma component of the picture portion; using the selected refinement parameter table to obtain one or more chroma refinement parameters for the block; and Refining the picture portion in the inter-component refinement mode, the inter-component refinement mode being applied block-by-block of the picture portion based on the obtained one or more chroma refinement parameters. one or more processors configured to perform Including, equipment.
19. The method of claim 14 , wherein a given value of the indication enables the inter-component refinement mode to be enabled for the block.
20. An apparatus as described in claim 13 or 18, wherein the one or more chroma improvement parameters are pivot points of points of a piecewise linear model.
21. The apparatus of claim 20, comprising constructing a lookup table using the one or more chroma improvement parameters.
22. The device described in claim 21, wherein the reconstructed version of the chroma components of the picture portion is improved using the constructed lookup table.
23. The device described in claim 13 or 18, wherein the multiple improvement parameter tables include multiple pairs of chroma improvement tables.
24. The apparatus of claim 13, further comprising decoding the plurality of improvement parameter tables from a slice header of the picture portion, the plurality of improvement parameter tables including one or more improvement parameters for improving the one or more chroma improvement parameters.
25. The apparatus of claim 18, further comprising encoding the plurality of improvement parameter tables from a slice header of the picture portion, the plurality of improvement parameter tables including one or more improvement parameters for improving the one or more chroma improvement parameters.
26. The apparatus described in claim 24 or 25, wherein the one or more improvement parameters include an index for updating an index of the one or more chroma improvement parameters and an improvement value for updating a value of the one or more chroma improvement parameters.
27. An apparatus as described in claim 13 or 18, wherein enabling or disabling the inter-component refinement mode for the block is estimated from the SAO parameters of the block.
Citation Information
Patent Citations
Filtering device, decoding apparatus and encoding apparatus
JP2013223050A
Crossplane filtering for chroma signal enhancement in video coding.
JP2015531569A
Coding data
JP2017118572A
Image encoder, image encoding method, image encoding program, image decoder, image decoding method, image decoding program, and image transmission system
JP2017188765A
Apparatus and method for image encoding, apparatus and method for image decoding, and image transmission system
US20170289541A1