Method and device for predictive picture encoding and decoding
Adaptive quantization of transformed prediction residuals addresses inefficiencies in transitioning to wider color gamut and dynamic range formats, enhancing coding performance by optimizing bitstream efficiency and reducing distortion.
Patent Information
- Application Number
- JP2023003422
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-04-07
- Filing Date
- 2023-01-12
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2038-03-15
AI Technical Summary
Existing video encoding technologies face inefficiencies when transitioning from BT.709 to wider color gamut and higher dynamic range formats like BT.2100, leading to signal distortion and loss of coding performance due to fixed-point precision mapping and inverse mapping processes.
Adapt the quantization step for transformed prediction residuals based on predicted sample values, deriving a mapping function to optimize bitstream efficiency by reducing bit cost or enhancing reconstruction quality.
Improves coding performance by minimizing bit cost or maximizing reconstruction quality through adaptive quantization, reducing distortion in video encoding and decoding processes.
Smart Images

Figure 0007812814000002 
Figure 0007812814000003 
Figure 0007812814000004
Abstract
Description
[Technical Field]
[0001] 1.Technical Field The present principles relate generally to methods and devices for picture encoding and decoding, and more particularly to methods and devices for encoding and decoding picture blocks. [Background technology]
[0002] 2.Background technology New generations of video formats include wider color gamuts, higher frame rates, and higher dynamic ranges. New standards have been created to support this type of content. For example, ITU-R Recommendation BT-2020 defines a format that includes primaries outside the color gamut of the currently used BT-709. ITU-R Recommendation BT-2100 defines a format that includes transfer functions that allow for an expansion of the dynamic range of content relative to BT.709. The primaries for BT-2100 are the same as those for BT-2020.
[0003] The use of BT.709 or BT.2100 containers results in significantly different codeword distributions. Most coding tools developed to date focus on SDR signals using BT.709 containers. Moving to wider containers, such as BT.2100, may require container adaptation or changes in codec design. Therefore, sample values need to be "reshaped" or mapped before coding to modify them in the new container to better match the properties expected by current codecs and encoders, such as HEVC. Summary of the Invention
[0004] It is known to perform a mapping / reshaping of samples represented in a given container (e.g., BT.2100) before encoding in order to obtain a sample distribution similar to that of the initial input samples (e.g., BT.709). The inverse mapping is applied to the decoded samples. The mapping before encoding and the inverse mapping after decoding cause distortion of the signal. In fact, both the mapping and inverse mapping processes are applied with fixed-point precision, which results in information loss. This distortion accumulates together with the distortion of the coding process, resulting in a loss of coding performance.
[0005] Instead of "reshaping" the sample values before coding, an alternative approach for processing the new container is to modify the quantization step for quantizing the coefficients of the transformed prediction residual. For this purpose, it is known to adapt the quantization step applied to the coefficients resulting from a transform (e.g., DCT) of the prediction residual samples for a given sample block based on values inferred from the predicted, original, or reconstructed samples of this block. Adapting the quantization step for each block can be inefficient, especially in cases where the block contains samples with many different values (e.g., bright and dark samples).
[0006] 3. Brief Overview 1. A method for coding a picture block, comprising the steps of: - obtaining a forecast; - determining mapped residual values from the source values and from the predicted values of the sample in response to a mapping function; - Encoding the mapped residual values and embedding them in the bitstream; wherein the mapping function is derived to obtain at least one of a reduced bit cost of the bitstream for a given reconstruction quality or an increased reconstruction quality for a given bit cost of the bitstream.
[0007] 1. A device for encoding picture blocks, comprising: - means for obtaining a prediction for at least one sample of the block and for one current component; - means for determining mapped residual values from the source values and from the predicted values of the sample in response to a mapping function; - means for encoding the mapped residual values and embedding them in the bitstream; wherein the mapping function is derived to obtain at least one of a reduced bit cost of the bitstream for a given quality of reconstruction or an increased quality of reconstruction for a given bit cost of the bitstream.
[0008] In a variant, an encoding device including a communication interface configured to access picture blocks and at least one processor, wherein the at least one processor: - obtaining a prediction for at least one sample of the accessed block and for one current component; - determining mapped residual values from the source values and from the predicted values of the sample in response to a mapping function; - Encoding the mapped residual values and embedding them in the bitstream; and wherein the mapping function is derived to obtain at least one of a reduced bit cost of the bitstream for a given quality of reconstruction or an increased quality of reconstruction for a given bit cost of the bitstream.
[0009] A bitstream representing a picture block, - coded data representing mapped residual values, the mapped residual values being obtained for at least one sample of the block and for one current component from source values and predicted values of the samples in response to a mapping function, the mapping function being derived to obtain at least one of a reduction in bit cost of the bitstream for a given reconstruction quality or an increase in reconstruction quality for a given bit cost of the bitstream; - coded data representing the mapping function and A bitstream is disclosed, comprising:
[0010] In a variant, a non-transitory processor-readable medium having stored thereon a bitstream representing picture blocks, the bitstream comprising: - coded data representing mapped residual values, the mapped residual values being obtained for at least one sample of the block and for one current component from source values and predicted values of the samples in response to a mapping function, the mapping function being derived to obtain at least one of a reduction in bit cost of the bitstream for a given reconstruction quality or an increase in reconstruction quality for a given bit cost of the bitstream; - coded data representing the mapping function and A non-transitory processor-readable medium is disclosed, including:
[0011] - transmitting coded data representing mapped residual values, the mapped residual values being obtained for at least one sample and one current component of the picture block from source values of the samples and from predicted values in response to a mapping function, the mapping function being derived to obtain at least one of a reduced bit cost of the bitstream for a given reconstruction quality or an increased reconstruction quality for a given bit cost of the bitstream; - transmitting coded data representing a mapping function; A transmission method is disclosed, including:
[0012] - means for transmitting coded data representing mapped residual values, the mapped residual values being obtained for at least one sample and one current component of the picture block from source values of the samples and from predicted values in response to a mapping function, the mapping function being derived to obtain at least one of a reduced bit cost of the bitstream for a given reconstruction quality or an increased reconstruction quality for a given bit cost of the bitstream; - means for transmitting coded data representing the mapping function; A transmitting device is disclosed, including:
[0013] 1. A transmitting device including a communication interface configured to access picture blocks and at least one processor, the at least one processor comprising: - transmitting coded data representing mapped residual values, the mapped residual values being obtained for at least one sample of the block and for one current component from source values and predicted values of the samples in response to a mapping function, the mapping function being derived to obtain at least one of a reduced bit cost of the bitstream for a given reconstruction quality or an increased reconstruction quality for a given bit cost of the bitstream; - transmitting coded data representing a mapping function; A transmitting device is disclosed that is configured to:
[0014] The following embodiments apply to the encoding method, encoding device, bitstream, processor-readable medium, transmitting method and transmitting device disclosed above.
[0015] In a first specific and non-limiting embodiment, determining the mapped residual value comprises: - mapping the source values of the samples using a mapping function; - mapping the predicted values of the samples using a mapping function; - determining a mapped residual value by subtracting the mapped predicted value from the mapped component value; Includes:
[0016] In a second specific and non-limiting embodiment, determining the mapped residual value comprises: - determining an intermediate residual value by subtracting the predicted value from the source value of the sample; - mapping intermediate residual values in response to a mapping function according to the predicted values; Includes:
[0017] In a third specific and non-limiting embodiment, mapping the intermediate residual values in response to the mapping function as a function of the predicted value includes multiplying the intermediate residual values by a scaling factor, where the scaling factor depends on the predicted value of the sample.
[0018] In a fourth specific and non-limiting embodiment, mapping the intermediate residual values in response to the mapping function according to the predicted values includes multiplying the intermediate residual values by a scaling factor, where the scaling factor depends on a predicted value obtained for another component of the sample, the other component being different from the current component.
[0019] 1. A method for decoding a picture block, comprising the steps of: - obtaining a forecast; - decoding residual values for the samples; - determining reconstruction values for the samples from the decoded residual values and from the prediction values in response to a mapping function; wherein the mapping function is derived to obtain at least one of a reduced bit cost of the bitstream for a given reconstruction quality or an increased reconstruction quality for a given bit cost of the bitstream.
[0020] 1. A device for decoding picture blocks, comprising: - means for obtaining a prediction for at least one sample of the block and for one current component; - means for decoding residual values for the samples; - means for determining a reconstruction value for the sample from the decoded residual values and from the prediction values in response to the mapping function; and wherein the mapping function is derived to obtain at least one of a reduced bit cost of the bitstream for a given quality of reconstruction or an increased quality of reconstruction for a given bit cost of the bitstream.
[0021] In a variant, a decoding device is provided, comprising a communication interface configured to access a bitstream and at least one process decoding or - obtaining a prediction for at least one sample of the block and for one current component; - decoding residual values for the samples from the accessed bitstream; - determining reconstruction values for the samples from the decoded residual values and from the prediction values in response to a mapping function; and wherein the mapping function is derived to obtain at least one of a reduced bit cost of the bitstream for a given quality of reconstruction or an increased quality of reconstruction for a given bit cost of the bitstream.
[0022] The following embodiments apply to the decoding method and decoding device disclosed above.
[0023] In a first specific and non-limiting embodiment, determining the reconstruction value for the sample comprises: - mapping the predicted values of the samples using a mapping function; - mapping the decoded residual values using an inverse of the mapping function; - determining a reconstruction value by adding the mapped prediction value to the mapped decoded residual value; Includes:
[0024] In a second specific and non-limiting embodiment, determining the reconstruction value for the sample comprises: - mapping the decoded residual values using the inverse of a mapping function as a function of the prediction; - determining the reconstruction value by adding the prediction value to the mapped decoded residual value; Includes:
[0025] In a third specific and non-limiting embodiment, mapping the decoded residual values using the inverse of the mapping function as a function of the predicted value comprises multiplying the decoded residual values by a scaling factor, where the scaling factor depends on the predicted value of the sample.
[0026] In a fourth specific and non-limiting embodiment, mapping the decoded residual value using the inverse of the mapping function as a function of the prediction value comprises multiplying the decoded residual value by a scaling factor, where the scaling factor depends on a prediction value obtained for another component of the sample, the other component being different from the current component. [Brief explanation of the drawings]
[0027] 4. A brief outline of the drawing [Figure 1] 1 illustrates an exemplary architecture of a transmitter configured to encode and embed pictures into a bitstream, according to a particular and non-limiting embodiment. [Figure 2] 1 illustrates an exemplary video encoder (eg, an HEVC video encoder) adapted to perform an encoding method in accordance with present principles. [Figure 3] 1 illustrates an exemplary architecture of a receiver configured to decode pictures from a bitstream to obtain decoded pictures, according to a particular and non-limiting embodiment. [Figure 4] 1 shows a block diagram of an exemplary video decoder (eg, an HEVC video decoder) adapted to perform a decoding method in accordance with the present principles. [Figure 5A] 1 depicts a flowchart of a method for encoding and embedding picture blocks into a bitstream, according to various embodiments. [Figure 5B] 1 depicts a flowchart of a method for decoding picture blocks from a bitstream, according to various embodiments. [Figure 6A] 1 depicts a flowchart of a method for encoding and embedding picture blocks into a bitstream, according to various embodiments. [Figure 6B] 1 depicts a flowchart of a method for decoding picture blocks from a bitstream, according to various embodiments. [Figure 7] Describe the mapping function fmap and its inverse function invfmap. [Figure 8A] 1 depicts a flowchart of a method for encoding and embedding picture blocks into a bitstream, according to various embodiments. [Figure 8B] 1 depicts a flowchart of a method for decoding picture blocks from a bitstream, according to various embodiments. [Figure 9] The derivative f'map of the mapping function fmap and the function 1 / f'map are depicted. [Figure 10A] 1 depicts a flowchart of a method for encoding and embedding picture blocks into a bitstream, according to various embodiments. [Figure 10B]1 depicts a flowchart of a method for decoding picture blocks from a bitstream, according to various embodiments. [Figure 11A] 1 depicts a flowchart of a method for encoding and embedding picture blocks into a bitstream, according to various embodiments. [Figure 11B] 1 depicts a flowchart of a method for decoding picture blocks from a bitstream, according to various embodiments. [Figure 12] 10 shows the full or limited range mapping functions constructed from the dQP table. DETAILED DESCRIPTION OF THE INVENTION
[0028] 5. Detailed Description It should be understood that the figures and descriptions are simplified to show elements relevant to a clear understanding of the present principles, while for clarity excluding many other elements found in a typical encoding and / or decoding device. It should be understood that, as used herein, the terms "first" and "second" may be used to describe various elements, and that these elements should not be limited by these terms. These terms are used only to distinguish one element from another.
[0029] A picture is an array of luma samples in monochrome format, or one array of luma samples and two corresponding arrays of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats. In general, a "block" addresses a specific area of the sample array (e.g., luma Y), and a "unit" includes an array block of all color components (luma Y and, possibly, chroma Cb and chroma Cr). A slice is an integer number of basic coding units, such as an HEVC coding tree unit or an H.264 macroblock unit. A slice can consist of a complete picture or a portion thereof. Each slice can include one or more slice segments.
[0030] In the following, the terms "reconstructed" and "decoded" can be used interchangeably. Typically (but not necessarily), "reconstructed" is used on the encoder side, and "decoded" is used on the decoder side. It should be noted that the terms "decoded" or "reconstructed" can mean that the bitstream is partially "decoded" or "reconstructed" (e.g., the signal obtained after deblocking filtering, but before SAO filtering), and that the reconstructed samples can differ from the final decoded output used for display. The terms "image," "picture," and "frame" can also be used interchangeably. The terms "sample" and "pixel" can also be used interchangeably.
[0031] Various embodiments are described with respect to the HEVC standard. However, the present principles are not limited to HEVC and may be applied to other standards, recommendations, and their extensions, including, for example, HEVC or HEVC extensions (such as Format Range (RExt), Scalability (SHVC), Multiview (MV-HEVC) extensions, and H.266). Various embodiments are described with respect to encoding / decoding of slices. Various embodiments may be applied to encoding / decoding of entire pictures or entire sequences of pictures.
[0032] References to "one embodiment" or "embodiment" of the present principles, and other variations thereof, mean that a particular feature, structure, characteristic, etc. described in connection with the embodiment is included in at least one embodiment. Thus, appearances of "one embodiment," "in an embodiment," "in one implementation," or "in an implementation," and any other variations thereof, appearing in various places throughout this specification are not necessarily all referring to the same embodiment.
[0033] It should be understood that the use of any of the following terms " / ," "and / or," and "at least one of" is intended to encompass, for example, in the cases of "A / B," "A and / or B," and "at least one of A or B," the selection of only the first listed option (A), the selection of only the second listed option (B), or the selection of both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, or C," such a statement is intended to encompass the selection of only the first listed option (A), the selection of only the second listed option (B), the selection of only the third listed option (C), the selection of only the first and second listed options (A and B), the selection of only the first and third listed options (A and C), the selection of only the second and third listed options (B and C), or the selection of all three options (A, B, and C). This can be expanded as many times as the number of items listed, as would be readily apparent to one of ordinary skill in this and related arts.
[0034] Various methods have been described above, each of which includes one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for the correct operation of the method, the order and / or use of specific steps and / or actions may be modified or combined.
[0035] FIG. 1 illustrates an exemplary architecture of a transmitter 1000 configured to encode and embed pictures into a bitstream, according to a particular and non-limiting embodiment.
[0036] The transmitter 1000 includes one or more processors 1005, which may include, for example, a CPU, GPU, and / or DSP (an English acronym for digital signal processor), along with internal memory 1030 (e.g., RAM, ROM, and / or EPROM). The transmitter 1000 includes one or more communication interfaces 1010 (e.g., a keyboard, mouse, touchpad, webcam), each adapted to display output information and / or allow a user to input commands and / or data, and a power supply 1020, which may be external to the transmitter 1000. The transmitter 1000 may also include one or more network interfaces (not shown). The encoder module 1040 represents a module that may be included in a device to perform encoding functions. Additionally, the encoder module 1040 may be implemented as a separate element of the transmitter 1000 or incorporated within the processor 1005, as a combination of hardware and software as known to those skilled in the art.
[0037] The picture can be obtained from sources, including but not limited to: - Local memory (e.g. video memory, RAM, flash memory, hard disk), - storage device interfaces (e.g. interfaces to mass storage devices, ROMs, optical disks or magnetic supports), a communication interface (e.g., a wired interface (e.g., a bus interface, a wide area network interface, a local area network interface) or a wireless interface (e.g., an IEEE 802.11 interface or a Bluetooth interface)), and - Picture capture circuitry (e.g., a sensor, such as a CCD (or Charge Coupled Device) or CMOS (or Complementary Metal Oxide Semiconductor)) It could be.
[0038] According to different embodiments, the bitstream can be transmitted to a destination. By way of example, the bitstream is stored in a remote or local memory (for example, video memory or RAM, hard disk). In a variant, the bitstream is sent to a storage interface (for example, an interface with a mass storage device, a ROM, a flash memory, an optical disk or a magnetic support) and / or transmitted over a communication interface (for example, an interface with a point-to-point link, a communication bus, a point-to-multipoint link or a broadcast network).
[0039] According to an exemplary and non-limiting embodiment, the Transmitter 1000 further includes a computer program stored in the memory 1030. The computer program includes instructions that, when executed by the Transmitter 1000 (specifically by the processor 1005), enable the Transmitter 1000 to perform the encoding method described with reference to Figures 5A, 6A, 8A, 10A, and 11A. According to a variant, the computer program is stored on a non-transitory digital data support external to the Transmitter 1000 (e.g., on an external storage medium such as a HDD, a CD-ROM, a DVD, a read-only and / or DVD drive, and / or a DVD read / write drive), all of which are known in the art. The Transmitter 1000 therefore includes a mechanism for reading the computer program. Furthermore, the Transmitter 1000 can access one or more Universal Serial Bus (USB) type storage devices (e.g., "memory sticks") through a corresponding USB port (not shown).
[0040] According to exemplary and non-limiting embodiments, the transmitter 1000 may include, but is not limited to: - mobile devices, - communication devices, - gaming devices, - a tablet (or tablet computer), - laptop, - Still picture camera, - video camera, - coding chips or coding devices / apparatus, - a still picture server, and - Video servers (e.g. broadcast servers, video-on-demand servers or web servers) It could be.
[0041] 2 shows an exemplary video encoder 100 (e.g., an HEVC video encoder) adapted to perform an encoding method according to one of the embodiments of FIGS. 5A, 6A, 8A, 10A, and 11 A. The encoder 100 is an example of a transmitter 1000 or a part of such a transmitter 1000.
[0042] For coding, a picture is typically partitioned into basic coding units (e.g., coding tree units (CTUs) in HEVC or macroblocks in H.264). A set of possibly contiguous basic coding units is grouped into a slice. A basic coding unit includes basic coding blocks of all color components. In HEVC, the minimum CTB size of 16x16 corresponds to the macroblock size as used in previous video coding standards. Although the terms CTU and CTB are used herein to describe encoding / decoding methods and encoding / decoding apparatus, it will be understood that these methods and apparatus should not be limited by these particular terms and may be expressed in other terms (e.g., macroblocks) in other standards, such as H.264.
[0043] In HEVC, the CTB is the root of a quadtree partitioned into coded blocks (CBs), which are partitioned into one or more predictive blocks (PBs) and form the root of a quadtree partitioned into transform blocks (TBs). Corresponding to coded blocks, predictive blocks, and transform blocks, coded units (CUs) include a tree-structured set of predictive units (PUs) and transform units (TUs), where a PU includes prediction information for all color components and a TU includes a residual coding syntax structure for each color component. The sizes of the CB, PB, and TB of the luma component apply to the corresponding CU, PU, and TU. In this application, the term "block" or "picture block" can be used to refer to any one of a CTU, CU, PU, TU, CB, PB, and TB. In addition, the term "block" or "picture block" can be used to refer to a macroblock, partition, and sub-block as specified in H.264 / AVC or other video coding standards, or more generally to an array of samples of various sizes.
[0044] In the exemplary encoder 100, a picture is encoded by the encoder elements as described below. The picture to be encoded is processed in units of CUs. Each CU is encoded using intra or inter mode. When a CU is encoded in intra mode, intra prediction (160) is performed. In inter mode, motion estimation (175) and compensation (170) are performed. The encoder decides (105) whether to use intra or inter mode to encode the CU and indicates the intra / inter decision with a prediction mode flag. The residual is calculated by subtracting (110) a prediction sample block (also known as a predictor) from the original picture block. The prediction sample block contains one prediction value for each sample of the block.
[0045] Intra modes, CUs are predicted from neighboring reconstructed samples within the same slice. A set of 35 intra prediction modes is available in HEVC, including DC, planar, and 33 angular prediction modes. The intra prediction reference is reconstructed from rows and columns adjacent to the current block. The reference spans more than twice the block size in the horizontal and vertical directions, using samples available from previously reconstructed blocks. When an angular prediction mode is used for intra prediction, the reference samples can be copied along the direction indicated by the angular prediction mode.
[0046] The luma intra-prediction modes applicable to the current block can be coded using two different options: If the applicable mode is included in a constructed list of three most probable modes (MPM), the mode is signaled by an index in the MPM list; otherwise, the mode is signaled by a fixed-length binarization of the mode index. The three most probable modes are derived from the intra-prediction modes of the neighboring blocks above and to the left.
[0047] For an inter CU, the corresponding coded block is further partitioned into one or more prediction blocks. Inter prediction is performed at the PB level, and the corresponding PU contains information about how inter prediction is performed.
[0048] Motion information (i.e., motion vectors and reference indices) can be signaled in two ways: "Advanced Motion Vector Prediction (AMVP)" and "Merge Mode." In AMVP, a video encoder or decoder assembles a candidate list based on motion vectors determined from already coded blocks. The video encoder then signals an index into the candidate list to identify a motion vector predictor (MVP) and signals a motion vector difference (MVD). At the decoder side, the motion vector (MV) is reconstructed as MVP+MVD.
[0049] In merge mode, a video encoder or decoder assembles a candidate list based on already coded blocks, and the video encoder signals an index to one of the candidates in the candidate list. At the decoder side, motion vectors and reference picture indices are reconstructed based on the signaled candidates.
[0050] In HEVC, the precision of motion information for motion compensation is 1 / 4 sample for the luma component and 1 / 8 sample for the chroma component. A 7-tap or 8-tap interpolation filter is used to interpolate the sample positions of the fractional samples. That is, 1 / 4, 1 / 2, and 3 / 4 of the total sample locations in both the horizontal and vertical directions can accommodate luma.
[0051] The residual is transformed (125) and quantized (130). The quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded (145) to output a bitstream. The encoder can also skip the transform and apply quantization directly to the untransformed residual signal on a 4x4 TU basis. The encoder can also avoid both the transform and quantization (i.e., the residual is directly coded without applying either the transform or quantization process). In direct PCM coding, no prediction is applied, and coded unit samples are directly coded and embedded in the bitstream.
[0052] The encoder includes a decoding loop, which decodes coded blocks to provide references for further prediction. Quantized transform coefficients are inverse quantized (140) and inverse transformed (150) to decode the residual. Combining the decoded residual with the predicted sample block (155) reconstructs a picture block. An in-loop filter (165) is applied to the reconstructed picture to perform, for example, deblocking / sample adaptive offset (SAO) filtering to reduce coding artifacts. The filtered picture can be stored in a reference picture buffer (180) and used as a reference for other pictures.
[0053] In HEVC, SAO filtering can be activated or deactivated at the picture level, slice level, and CTB level. Two SAO modes are specified: Edge Offset (EO) and Band Offset (BO). For EO, sample classification is based on the local directional structure of the picture to be filtered. For BO, sample classification is based on the sample value. Parameters for EO or BO can be explicitly coded or derived from neighborhoods. SAO can be applied to luma and chroma components, and the SAO modes are the same for Cb and Cr components. The SAO parameters (i.e., offset, SAO type EO, BO, and deactivation, class in the case of EO, band position in the case of BO) are configured separately for each color component.
[0054] FIG. 3 illustrates an exemplary architecture of a receiver 2000 configured to decode pictures from a bitstream to obtain decoded pictures, according to a particular and non-limiting embodiment.
[0055] The receiver 2000 includes one or more processors 2005, which may include, for example, a CPU, a GPU, and / or a DSP (an English acronym for digital signal processor), along with internal memory 2030 (e.g., RAM, ROM, and / or EPROM). The receiver 2000 includes one or more communication interfaces 2010 (e.g., a keyboard, a mouse, a touchpad, a webcam), each adapted to display output information and / or allow a user to input commands and / or data (e.g., decoded pictures), and a power supply 2020, which may be external to the receiver 2000. The receiver 2000 may also include one or more network interfaces (not shown). The decoder module 2040 represents a module that may be included in a device to perform decoding functions. Additionally, the decoder module 2040 may be implemented as a separate element of the receiver 2000 or incorporated within the processor 2005, as a combination of hardware and software as known to those skilled in the art.
[0056] The bitstream can be obtained from sources, including but not limited to: - Local memory (e.g. video memory, RAM, flash memory, hard disk), - storage device interfaces (e.g. interfaces to mass storage devices, ROMs, optical disks or magnetic supports), a communication interface (e.g., a wired interface (e.g., a bus interface, a wide area network interface, a local area network interface) or a wireless interface (e.g., an IEEE 802.11 interface or a Bluetooth interface)), and - Image capture circuitry (e.g., sensors such as CCD (or Charge Coupled Device) or CMOS (or Complementary Metal Oxide Semiconductor)) It could be.
[0057] According to different embodiments, the decoded pictures can be transmitted to a destination (e.g. a display device). By way of example, the decoded pictures are stored in a remote or local memory (e.g. a video memory or RAM, a hard disk). In a variant, the decoded pictures are sent to a storage interface (e.g. an interface to a mass storage device, a ROM, a flash memory, an optical disk or a magnetic support) and / or transmitted over a communication interface (e.g. an interface to a point-to-point link, a communication bus, a point-to-multipoint link or a broadcast network).
[0058] According to a specific and non-limiting embodiment, the receiver 2000 further includes a computer program stored in the memory 2030. The computer program includes instructions that, when executed by the receiver 2000 (specifically by the processor 2005), enable the receiver to perform the decoding method described with reference to Figures 5B, 6B, 8B, 10B, and 11B. According to a variant, the computer program is stored on a non-transitory digital data support external to the receiver 2000 (e.g., on an external storage medium such as a HDD, a CD-ROM, a DVD, a read-only and / or DVD drive, and / or a DVD read / write drive), all of which are known in the art. The receiver 2000 therefore includes a mechanism for reading the computer program. Furthermore, the receiver 2000 can access one or more Universal Serial Bus (USB) type storage devices (e.g., "memory sticks") through corresponding USB ports (not shown).
[0059] According to an exemplary and non-limiting embodiment, the receiver 2000 may include, but is not limited to: - mobile devices, - communication devices, - gaming devices, - set-top boxes, - TV sets, - a tablet (or tablet computer), - laptop, - Video players (e.g. Blu-ray players, DVD players), - Display, and - Decryption chip or decryption device / apparatus It could be.
[0060] Figure 4 shows a block diagram of an exemplary video decoder 200 (e.g., an HEVC video decoder) adapted to perform a decoding method according to an embodiment of Figures 5B, 6B, 8B, 10B, and 11B. The video decoder 200 is an example of a receiver 2000 or a portion of such a receiver 2000. In the exemplary decoder 200, the bitstream is decoded by decoder elements as described below. The video decoder 200 generally performs a decoding pass that is the inverse of the encoding pass as described in Figure 2, which performs video decoding as part of encoding the video data.
[0061] Specifically, the decoder's input includes a video bitstream, which may be generated by video encoder 100. The bitstream is first entropy decoded (230) to obtain transform coefficients, motion vectors, and other coded information. The transform coefficients are inversely quantized (240) and inversely transformed (250) to decode the residual. The decoded residual is then combined (255) with a prediction sample block (also known as a predictor) to obtain a decoded / reconstructed picture block. The prediction sample block may result from intra prediction (260) or motion-compensated prediction (i.e., inter prediction) (270). As described above, AMVP and merge mode techniques may be used during motion compensation, which may use an interpolation filter to calculate interpolated values for sub-integer samples of the reference block. An in-loop filter (265) is applied to the reconstructed picture. The in-loop filter may include a deblocking filter and an SAO filter. The filtered picture is stored in a reference picture buffer (280).
[0062] 5A shows a flowchart of a method for encoding and embedding picture blocks into a bitstream according to the present principles. Mapping is applied in a coding loop to obtain pixel-level mapped residual samples. In contrast to the prior art, the input samples of the encoding method are not modified by the mapping. On the decoder side, the output samples from the decoder are not modified by the inverse mapping. Mapping can be applied to one or several components of a picture. For example, mapping can be applied to only the luma component, only the chroma component, or both the luma and chroma components.
[0063] The method begins at step S100. In step S110, the transmitter 1000 (e.g., the encoder 100) accesses a block of a picture slice. In step S120, the transmitter obtains a predicted value Pred(x,y) of the source value Orig(x,y) for at least one sample of the accessed block and for at least one component (e.g., for luma), where (x,y) are the spatial coordinates of the sample in the picture. The predicted value is obtained (i.e., is typically determined) depending on the prediction mode (intra / inter mode) selected for the block.
[0064] In step S130, the transmitter calculates a mapping function f map The method determines mapped residual values from the sample source values Orig(x,y) and from the predicted values Pred(x,y) in response to (). The mapping function is defined or derived to obtain a coding gain, i.e., a reduction in the bit-cost (i.e., number of bits) of the bitstream for a given visual or objective reconstruction quality, or an increase in the visual or objective reconstruction quality for a given bit-cost. When a block, picture, or picture sequence is coded and embedded into a bitstream of a given size (i.e., a given number of bits), the receiver-side reconstruction quality of the block, picture, or picture sequence depends on this size. On the other hand, when a block, picture, or picture sequence is coded with a given reconstruction quality, the size of the bitstream depends on this reconstruction quality.
[0065] In most cases, distortion, which represents the quality of the reconstruction, is defined as the expected squared difference between the input and output signals (i.e., the mean squared error). However, since most lossy compression techniques operate on data perceived by human consumers (viewers of pictures and videos), the distortion measure can preferably be modeled based on human perception, possibly aesthetically.
[0066] For example, the mapping function can be derived by one of the following approaches: The mapping function is derived in such a way that the amplitude of the residual value increases more for component values with large amplitude values than for component values with small amplitude values, as depicted in FIG. - To obtain improved perceptual or objective coding performance, a predefined encoder quantization adjustment table deltaQP or quantization adjustment function dQP(Y), where Y is the video signal luma, can be derived or adjusted. From deltaQP or dQP(Y), a scaling function can be derived as follows: sc(Y)=2^(-dQP(Y) / 6), where ^ is the exponentiation operator. The scaling function can be used in a mapping function and can amount to a multiplication of the residual with a scaling value derived from the scaling function. In a variant, the mapping function can be derived by considering that this scaling function is the derivative of the mapping function applied to the residual. In step S130, to map the residual, the pre-encoder function Map(Y) (Y is the luma video signal) or a derivative of Map(Y) (which is a scaling function) is applied to a mapping function f map The pre-encoder function Map(Y) is derived in such a way that the original samples of the signal (once mapped by this pre-encoder function Map(Y)) are better distributed over the entire codeword range (e.g., thanks to histogram equalization).
[0067] In addition to the three techniques mentioned above, other techniques can be used to derive the mapping function if the mapping of residual values improves compression performance.
[0068] Steps S110 and S120 may be repeated for each sample of the accessed block to obtain a block of mapped residual values.
[0069] In step S140, the transmitter encodes the mapped residual values, which typically, but not necessarily, includes converting the residuals into transform coefficients, quantizing the coefficients with a quantization step size QP to obtain quantized coefficients, and entropy coding the quantized coefficients for embedding in the bitstream.
[0070] The method ends in step S180.
[0071] FIG. 5B shows a flowchart of a method for decoding picture blocks of a bitstream corresponding to the encoding method of FIG. 5A.
[0072] The method begins at step S200. In step S210, the receiver 2000 (eg, decoder 200) accesses the bitstream.
[0073] In step S220, the receiver obtains a prediction value Pred(x,y) for at least one sample for at least one component (e.g., luma), where (x,y) are the spatial coordinates of the sample in the picture. The prediction value is obtained according to the prediction mode (intra / inter mode) selected for the block.
[0074] In step S230, the receiver decodes the residual value Res(x,y) for the sample to be decoded. The residual value Res(x,y) is a decoded version of the mapped residual value that was coded in step S140 of Figure 5A. Decoding typically, but not necessarily, involves entropy decoding a portion of the bitstream representing the block to obtain a block of transform coefficients, and inverse quantizing and inverse transforming the block of transform coefficients to obtain a block of residuals.
[0075] In step S240, the transmitter applies to the samples the mapping function f used by the encoding method in step S130. map A mapping function invf that is the inverse of () mapThe reconstructed sample values are determined from the decoded residual values and from the predicted values in response to (). Steps S220 to S240 may be repeated for each sample of the accessed block.
[0076] The method ends in step S280.
[0077] FIG. 6A shows a flowchart of a method for encoding and embedding picture blocks into a bitstream according to a first particular and non-limiting embodiment.
[0078] The method begins at step S100. In step S110, the transmitter 1000 (e.g., the encoder 100) accesses a block of a picture slice. In step S120, the transmitter obtains a predicted value Pred(x,y) of the value Orig(x,y) for at least one sample of the accessed block for at least one component (e.g., for luma), where (x,y) are the spatial coordinates of the sample in the picture. The predicted value is obtained depending on the prediction mode (intra / inter mode) selected for the block.
[0079] In step S130, the transmitter calculates a mapping function f map The method determines mapped residual values from the source values Orig(x,y) and from the predicted values Pred(x,y) of the samples in response to (). The mapping function is defined or derived to obtain a coding gain, i.e., a reduction in bit rate for a given visual or objective quality or an increase in visual or objective quality for a given bit rate. The mapping function can be derived by one of the methods disclosed with reference to FIG. 5A. Steps S110 to S130 can be repeated for each sample of the accessed block to obtain a block of mapped residual values. In the first embodiment, Res map The mapped residual, denoted by (x,y), is map (Orig(x,y))-f map Equals (Pred(x,y)).
[0080] In step S140, the transmitter encodes the mapped residual values, which typically, but not necessarily, includes converting the residuals into transform coefficients, quantizing the coefficients with a quantization step size QP to obtain quantized coefficients, and entropy coding the quantized coefficients for embedding in the bitstream.
[0081] The method ends in step S180.
[0082] FIG. 6B represents a flowchart of a method for decoding picture blocks of a bitstream corresponding to the embodiment of the encoding method according to FIG. 6A, according to a first particular and non-limiting embodiment.
[0083] The method begins at step S200. In step S210, the receiver 2000 (eg, decoder 200) accesses the bitstream.
[0084] In step S220, the receiver obtains a prediction value Pred(x,y) for at least one sample for at least one component (e.g., luma), where (x,y) are the spatial coordinates of the sample in the picture. The prediction value is obtained according to the prediction mode (intra / inter mode) selected for the block.
[0085] In step S230, the receiver decodes the residual value Res(x,y) for the sample to be decoded. The residual value Res(x,y) is a decoded version of the mapped residual value that was coded in step S140 of Figure 6A. Decoding typically, but not necessarily, involves entropy decoding a portion of the bitstream representing the block to obtain a block of transform coefficients, and inverse quantizing and inverse transforming the block of transform coefficients to obtain a block of residuals.
[0086] In step S240, the transmitter applies to the samples the mapping function f used by the encoding method in step S130. map () and its inverse function invf map () and from the decoded residual values Res(x,y) and from the predicted values Pred(x,y). Steps S220 to S240 can be repeated for each sample of the accessed block to obtain a reconstructed block. In the first embodiment, the reconstructed sample values denoted Dec(x,y) are calculated by the invf map (Res(x,y)+f map Equivalent to (Pred(x,y)).
[0087] The method ends in step S280.
[0088] FIG. 8A illustrates a flowchart of a method for encoding and embedding picture blocks into a bitstream according to a second specific and non-limiting embodiment.
[0089] The method begins at step S100. In step S110, the transmitter 1000 (e.g., the encoder 100) accesses a block of a picture slice. In step S120, the transmitter obtains a predicted value Pred(x,y) of the value Orig(x,y) for at least one sample of the accessed block for at least one component (e.g., for luma), where (x,y) are the spatial coordinates of the sample in the picture. The predicted value is obtained depending on the prediction mode (intra / inter mode) selected for the block.
[0090] In step S130, the transmitter calculates a mapping function g mapThe method determines mapped residual values from the source values Orig(x,y) and from the predicted values Pred(x,y) of the samples in response to (). The mapping function is defined or derived to obtain a coding gain, i.e., a reduction in bit rate for a given visual or objective quality or an increase in visual or objective quality for a given bit rate. The mapping function can be derived by one of the methods disclosed with reference to FIG. 5A. Steps S110 to S130 can be repeated for each sample of the accessed block to obtain a block of mapped residual values. In the second embodiment, Res map The mapped residual, denoted by (x,y), is map (Res usual (x,y),Pred(x,y)), where Res usual (x,y) = Orig(x,y) - Pred(x,y).
[0091] Function g map (p,v) and invg map One simple version of (p,v) can be derived from the first embodiment: for a predicted value p and a sample residual value v, g map (p,v) and invg map (p,v) can be constructed as follows:
[0092] In the first embodiment, Res remap (x,y)=f map (Orig(x,y))-f map (Pred(x,y)) When the signals Orig(x,y) and Pred(x,y) are close (which is expected if the prediction works well), we can consider Orig(x,y) = Pred(x,y) + ε (where ε is a very small amplitude). Considering the definition of the derivative of a function, f map (Orig(x,y))=f map (Pred(x,y)+ε)≒f map (Pred(x,y))+ε * f' map(Pred(x,y)) It can be considered that In the formula, f map corresponds to a 1D function (e.g., as defined in embodiment 1), and f' map is a function f map is the derivative of Next, Res map (x,y)=f map (Orig(x,y))-f map (Pred(x,y))≒ε * f' map (Pred(x,y)). By definition, ε=Orig(x,y)-Pred(x,y) is the normal prediction residual Res usual (x,y). Therefore, the following function g map (p,v) and invg map (p,v) can be used. g map (p,v)=f' map (p) * v invg map (p,v)=(1 / f' map (p)) * v
[0093] On the encoder side, the mapped residual is Res map (x,y)=f' map (Pred(x,y)) * Res usual (x,y) (eq.1) is derived as: At the decoder side, the reconstructed signal is Dec(x,y)=Pred(x,y)+1 / f' map (Pred(x,y)) * Res dec (x,y)) (eq.2) is derived as:
[0094] This means that the mapping is usually a simple scaling of the residual by a scaling factor that depends on the predicted value. Sometimes, on the encoder side, the scaling factor depends on the original value, not the predicted value. However, doing so creates a mismatch between the encoder and the decoder. It is also possible to use a filtered version of the prediction, for example by using a smoothing filter, to reduce the effect of quantization errors. For example, instead of using Pred(x,y) in equations 1 and 2, one can use the filtered version (Pred(x-1,y) / 4+Pred(x,y) / 2+Pred(x+1,y)) / 4).
[0095] Figure 9 shows the function f' map and (1 / f' map ) provides an example.
[0096] In step S140, the transmitter encodes the mapped residual values, which typically, but not necessarily, includes converting the residuals into transform coefficients, quantizing the coefficients with a quantization step size QP to obtain quantized coefficients, and entropy coding the quantized coefficients for embedding in the bitstream.
[0097] The method ends in step S180.
[0098] FIG. 8B shows a flowchart of a method for decoding picture blocks of a bitstream corresponding to the encoding method disclosed with respect to FIG. 8A.
[0099] The method begins at step S200. In step S210, the receiver 2000 (eg, decoder 200) accesses the bitstream.
[0100] In step S220, the receiver obtains a prediction value Pred(x,y) for at least one sample for at least one component (e.g., luma), where (x,y) are the spatial coordinates of the sample in the picture. The prediction value is obtained according to the prediction mode (intra / inter mode) selected for the block.
[0101] In step S230, the receiver decodes the residual value Res(x,y) for the sample to be decoded. The residual value Res(x,y) is a decoded version of the mapped residual value that was coded in step S140 of Figure 8A. Decoding typically, but not necessarily, involves entropy decoding a portion of the bitstream representing the block to obtain a block of transform coefficients, and inverse quantizing and inverse transforming the block of transform coefficients to obtain a block of residuals.
[0102] In step S240, the transmitter applies to the samples the mapping function g used by the encoding method in step S130 of FIG. 8A. map The mapping function invg, which is the inverse of () map In response to (), a reconstructed sample value Dec(x,y) is determined from the decoded residual value Res(x,y) and from the predicted value Pred(x,y). Steps S220 to S240 can be repeated for each sample of the accessed block to obtain a reconstructed block. In a second embodiment, the reconstructed sample value denoted Dec(x,y) is calculated as Pred(x,y)+invg map Equals (Res(x,y),Pred(x,y)).
[0103] The first embodiment is f map Functions and invf map While this embodiment requires applying both invg and map It allows to map the prediction residuals using a function.
[0104] The method ends in step S280.
[0105] 10A shows a flowchart of a method for encoding and embedding picture blocks into a bitstream according to a third specific and non-limiting embodiment. This embodiment is a generalization of the second embodiment. The function f map and invf map () is a scaling function, the scaling factor of which depends on the value of the predicted signal (or a filtered version of the predicted signal as mentioned earlier).
[0106] The method begins at step S100. In step S110, the transmitter 1000 (e.g., the encoder 100) accesses a block of a picture slice. In step S120, the transmitter obtains a predicted value Pred(x,y) of the value Orig(x,y) for at least one sample of the accessed block for at least one component (e.g., for luma), where (x,y) are the spatial coordinates of the sample in the picture. The predicted value is obtained depending on the prediction mode (intra / inter mode) selected for the block.
[0107] In step S130, the transmitter calculates a mapping function f map The method determines mapped residual values from the source values Orig(x,y) and from the predicted values Pred(x,y) of the samples in response to (). The mapping function is defined or derived to obtain a coding gain, i.e., a reduction in bit rate for a given visual or objective quality or an increase in visual or objective quality for a given bit rate. The mapping function can be derived by one of the methods disclosed with reference to FIG. 5A. Steps S110 to S130 can be repeated for each sample of the accessed block to obtain a block of mapped residual values. In the second embodiment, Res map The mapped residual, denoted by (x,y), is map (Pred(x,y)) * Res usual (x,y) where Resusual (x,y) = Orig(x,y) - Pred(x,y). This is a generalized version of (eq.1) and (eq.2). In a variant, the original value Orig(x,y) can be used instead of Pred(x,y). In this case, Res map (x,y) is f map (Orig(x,y)) * Res usual In another variant, a combination of Orig(x,y) and Pred(x,y), Comb(x,y), can be used (e.g., the average of these two values). In this latter case, Res map (x,y) is f map (Comb(x,y)) * Res usual Equal to (x,y).
[0108] In step S140, the transmitter encodes the mapped residual values, which typically, but not necessarily, includes converting the residuals into transform coefficients, quantizing the coefficients with a quantization step size QP to obtain quantized coefficients, and entropy coding the quantized coefficients for embedding in the bitstream.
[0109] The method ends in step S180.
[0110] FIG. 10B shows a flowchart of a method for decoding picture blocks from a bitstream corresponding to the encoding method disclosed with respect to FIG. 10A.
[0111] The method begins at step S200. In step S210, the receiver 2000 (eg, decoder 200) accesses the bitstream.
[0112] In step S220, the receiver obtains a prediction value Pred(x,y) for at least one sample for at least one component (e.g., luma), where (x,y) are the spatial coordinates of the sample in the picture. The prediction value is obtained according to the prediction mode (intra / inter mode) selected for the block.
[0113] In step S230, the receiver decodes the residual value Res(x,y) for the sample to be decoded. The residual value Res(x,y) is a decoded version of the mapped residual value that was coded in step S140 of Figure 10A. Decoding typically, but not necessarily, involves entropy decoding a portion of the bitstream representing the block to obtain a block of transform coefficients, and inverse quantizing and inverse transforming the block of transform coefficients to obtain a block of residuals.
[0114] In step S240, the receiver applies to the samples the mapping function invf used by the encoding method in step S130. map ()=1 / f map In response to (), a reconstructed sample value Dec(x,y) is determined from the decoded residual value Res(x,y) and from the prediction value Pred(x,y). Steps S220 to S240 can be repeated for each sample of the accessed block to obtain a reconstructed block. In a second embodiment, the reconstructed sample value denoted Dec(x,y) is calculated as Pred(x,y)+(1 / f map (Pred(x,y))) * Equal to Res(x,y).
[0115] This embodiment advantageously allows the inverse mapping to be performed at the decoder side by using simple multiplications, thereby limiting the additional complexity and allowing the exact mapping to be performed, with the rounding operation being applied at the very end of the process (when computing Dec(x,y)).
[0116] The method ends in step S280.
[0117] 11A shows a flowchart of a method for encoding and embedding picture blocks in a bitstream according to a fourth specific and non-limiting embodiment. In this embodiment, the mapping is a cross-component scaling. For example, the mapping is applied to the chroma component C (C is U (or Cb) or V (or Cr)) depending on the co-located luma component Y (or its filtered version). When the luma or chroma pictures are not of the same resolution (e.g., in the case of a 4:2:0 chroma format), the luma value can be taken after resampling or as one of the sample values of the luma picture with which the chroma sample is associated. For example, in the case of a 4:2:0 signal, for a position (x,y) in the picture, the luma value at position (2 * x,2 * y) can be considered.
[0118] The method begins at step S100. In step S110, the transmitter 1000 (e.g., the encoder 100) accesses a block of a picture slice. In step S120, the transmitter obtains a predicted value PredC(x,y) of a source value OrigC(x,y) for at least one sample of the accessed block for at least one component (e.g., for chroma C), where (x,y) are spatial coordinates of the sample in the picture. The transmitter further obtains a predicted value PredY(x,y) of a source value OrigY(x,y) for the same sample for at least another component (e.g., for luma Y). The predicted value is obtained depending on the prediction mode (intra / inter mode) selected for the block.
[0119] In step S130, the transmitter calculates a mapping function f mapThe method determines mapped residual values from the source values OrigC(x,y) of the samples and from the predicted values PredC(x,y) and PredY(x,y) in response to (). The mapping function is defined or derived to obtain a coding gain, i.e., a reduction in bit rate for a given visual or objective quality or an increase in visual or objective quality for a given bit rate. The mapping function can be derived by one of the methods disclosed with reference to FIG. 5A. Steps S110 to S130 can be repeated for each sample of the accessed block to obtain a block of mapped residual values. In the fourth embodiment, ResC map The mapped residual, denoted by (x,y), is map (PredY(x,y)) * ResC usual (x,y) where ResC usual (x,y) = OrigC(x,y) - PredC(x,y), where OrigC(x,y) is the value of the source sample of chrominance component C (to be coded) at position (x,y) of the picture, PredC(x,y) is the value of the predicted sample of chrominance component C, and ResC usual (x, y) is the value of the prediction residual sample of the chroma component C.
[0120] In step S140, the transmitter encodes the mapped residual values, which typically, but not necessarily, includes converting the residuals into transform coefficients, quantizing the coefficients with a quantization step size QP to obtain quantized coefficients, and entropy coding the quantized coefficients for embedding in the bitstream.
[0121] The method ends in step S180.
[0122] FIG. 11B shows a flowchart of a method for decoding picture blocks from a bitstream corresponding to the encoding method disclosed with respect to FIG. 11A.
[0123] The method begins at step S200. In step S210, the receiver 2000 (eg, decoder 200) accesses the bitstream.
[0124] In step S220, the receiver obtains, for at least one component (e.g., for chroma C), a predicted value PredC(x,y) of a source value OrigC(x,y) for at least one sample of the accessed block, where (x,y) are the spatial coordinates of the sample in the picture. The receiver further obtains, for at least another component (e.g., for luma Y), a predicted value PredY(x,y) of a source value OrigY(x,y) for the same sample (possibly downsampled if the luma and chroma pictures do not have the same resolution). The predicted value is obtained depending on the prediction mode (intra / inter mode) selected for the block.
[0125] In step S230, the receiver decodes the residual value ResC(x,y) for the sample to be decoded. The residual value ResC(x,y) is a decoded version of the mapped residual value that was coded in step S140 of Figure 11A. Decoding typically, but not necessarily, involves entropy decoding a portion of the bitstream representing the block to obtain a block of transform coefficients, and inverse quantizing and inverse transforming the block of transform coefficients to obtain a block of residuals.
[0126] In step S240, the receiver calculates f map where () is the mapping function 1 / f map In response to (), a reconstructed sample value DecC(x,y) is determined from the decoded residual value ResC(x,y) and from the prediction values PredC(x,y) and PredY(x,y). Steps S220 to S240 may be repeated for each sample of the accessed block to obtain a reconstructed chroma block. In the fourth embodiment, the reconstructed sample value, denoted DecC(x,y), is calculated as PredC(x,y)+(1 / f map(PredY(x,y))) * Equivalent to ResC(x,y).
[0127] This embodiment advantageously allows scaling the chroma components depending on the luma components at the decoder side, which generally improves visual quality thanks to finer control of chroma scaling for different luma intervals.
[0128] The method ends in step S280.
[0129] The third and fourth embodiments disclosed with respect to Figures 10A, 10B, 11A and 11B can advantageously be implemented in a fixed-point manner.
[0130] ResC usual Let be the prediction residual to be mapped at location (x,y). The scaling factor scal determines, for example, in the case of cross-component scaling, f map From the value of scal=f map (PredY) is derived.
[0131] On the decoder side, invScal=round(2^B÷scal) is used (possibly coded and embedded in the bitstream, as explained below), During the ceremony, ^ is the exponentiation operator, round(x) is the nearest integer value of x, B is the bit depth chosen to quantize the scaling factor (typically B=8 or 10 bits).
[0132] Mapped value ResC map Value to ResC usual The mapping is applied as follows: ResC map =(ResCusual * 2 B +sign(ResC usual ) * (invScal / 2)) / invScal (eq.3) In the formula, ResC usual (x,y)=OrigC(x,y)-PredC(x,y), and if x≧0, then sign(x) is equal to 1, otherwise it is equal to −1. All parameters in this equation are integers, and the " / " division is also applied in integers (the "÷" division is a floating point division). Then, the mapped value ResC map is encoded.
[0133] At the decoder side, the encoded mapped value ResC map is the value ResC map_dec The reverse mapped value ResC is invmap The decrypted value ResC map_dec The inverse mapping of is applied as follows: ResC invmap =(ResC map_dec * invScal+sign(ResC map_dec ) * 2 (B-1) ) / 2 B (eq.4) (ResCmap_dec * invScal+sign(ResCmap_dec) * 2(B-1)) / 2B this is, ResC invmap =(ResC map_dec * invScal+sign(ResC map_dec ) * 2 (B-1) )>>B (eq.5) is equal to.
[0134] Then, at location (x, y), the predicted values PredC and ResC invmap from, DecC=PredC+ResC invmap (eq.6) The reconstruction value DecC is derived as follows.
[0135] We can also combine these operations directly to avoid using the sign operator: Equations (eq.5) and (eq.6) are combined to give (eq.7): DecC=(PredC * 2 B +ResC map_dec * invScal+2 (B-1) )>>B (eq.7).
[0136] In HEVC, quantization is adjusted using a quantization parameter QP. From QP, a quantization step Qstep0 is derived, which is given by (K * 2^(QP / 6)), where K is a fixed parameter. When local QP correction dQP is used, the actual quantization step Qstep1 is (K * It can be approximated as 2^((QP+dQP) / 6)), i.e., (Qstep0 * 2^(dQP / 6)) The signal is divided by the quantization step. This means that for a given dQP, the corresponding scaling applied to the signal during quantization (derived from the inverse of the quantization step) is equal to 2^(-dQP / 6). For example, the following correspondence can be established for the dQP table:
[0137] [Table 1]
[0138] The scaling can be used, for example, in the scaling solution described in the third embodiment. The scaling can also be used to derive a mapping function as used in the first and second embodiments. In fact, this scaling corresponds to the derivative of the mapping function. Therefore, the mapping function can be modeled as a piecewise linear function, with each piece having a slope equal to the scaling corresponding to this piece. The dQP associated with each interval i A set of intervals [Y i ,Y i+1 −1] (i=0 to n, n is an integer), the mapping function can be defined as follows: Let i be the index of the interval containing Y (Y is [Y i ,Y i+1 -1]), f map (Y)=f map (Y i )+2^(-dQP i / 6) * (YY i ) This results in the functions shown in FIG. 12 (for full range (FR) or limited range (LR) signal representations) for the particular dQP table above.
[0139] Function f map or g map or their inverse functions invf map or invg map can be explicitly defined in the decoder (hence in the decoder specification) or signaled in the bitstream. Function f map , invf map , g map or invg map teeth, Lookup tables Piecewise Scalar Functions (PWS) Piecewise Linear Function (PWL) Piecewise Polynomial Functions (PWP) It can be implemented in the form of These functions can be coded in new structures such as SEI messages, sequence parameter sets (SPS), picture parameter sets (PPS), slice headers, coding tree unit (CTU) syntax, per tile, or adaptive picture sets (APS).
[0140] Implementations described herein may be implemented as, for example, a method or process, an apparatus, a software program, a data stream, or a signal. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method or a device), the implementation of the discussed features may also be implemented in other forms (e.g., a program). An apparatus may be implemented, for example, in appropriate hardware, software, and firmware. A method may be implemented, for example, in an apparatus (e.g., processor, etc.), which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end users.
[0141] Implementations of the various processes and features described herein may be embodied in a variety of different devices or applications (e.g., devices or applications, among others). Examples of such devices include encoders, decoders, post-processors that process output from decoders, pre-processors that provide input to encoders, video coders, video decoders, video codecs, web servers, set-top boxes, laptops, personal computers, mobile phones, PDAs, and other communication devices. It should be clear that the devices may be mobile, even installed in moving vehicles.
[0142] Additionally, methods may be implemented by instructions being executed by a processor, and such instructions (and / or data values produced by an implementation) may be stored on a processor-readable medium such as, for example, an integrated circuit, a software carrier, or other storage device (e.g., a hard disk, a compact disc (“CD”), an optical disc (e.g., a DVD, often referred to as a digital versatile disc or digital video disc), a random access memory (“RAM”), or a read-only memory (“ROM”)). The instructions may form an application program tangibly embodied on the processor-readable medium. The instructions may be, for example, in hardware, firmware, software, or a combination. The instructions may be found, for example, in an operating system, a separate application, or a combination of the two. A processor may thus be characterized as both, for example, a device configured to execute a process and a device that includes a processor-readable medium (e.g., a storage device) having instructions for executing a process. Furthermore, a processor-readable medium may store data values produced by an implementation in addition to or in place of instructions.
[0143] As will be apparent to those skilled in the art, implementations can generate various signals formatted to carry information that can be stored or transmitted, for example. Information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry rules for writing or reading syntax of the described embodiments as data, or to carry the actual syntax values written by the described embodiments as data. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using a high frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.
[0144] A number of implementations have been described. Nevertheless, it will be understood that various modifications may be made. For example, elements of different implementations may be combined, supplemented, modified, or removed to produce other implementations. Additionally, those skilled in the art will understand that other structures and processes may be substituted for those disclosed, with the resulting implementations performing at least substantially the same function, in at least substantially the same way, to achieve at least substantially the same results as the disclosed implementations. Accordingly, these and other implementations are contemplated by this application.
Claims
1. A method of encoding, comprising: - obtaining a luma prediction value of a luma source value of a luma component of a sample of a block of a picture and a chroma prediction value of a chroma source value of a chroma component of said sample, said luma prediction value and said chroma prediction value being obtained by spatial prediction or temporal prediction of said block; determining mapped luminance residual values from the luminance source values of the luminance component and from the luminance prediction values using a mapping function, - mapping said luminance prediction value with said mapping function to obtain a mapped luminance prediction value; - mapping said luminance source values using said mapping function to obtain said mapped luminance source values; determining a mapped luma residual value, the mapped luma residual value being equal to the difference of the mapped luma prediction value from the mapped luma source value; determining mapped chrominance residual values from the chrominance source values of the chrominance components and from the chrominance predicted values using values representing the mapped luma predicted values; encoding and embedding said mapped luma residual values and said mapped chroma residual values into a bitstream; A method comprising:
2. 1. A method of decoding, comprising: obtaining luminance predictions of the luminance component of the samples of the picture block; obtaining color difference predictions for the color difference components of said samples; - decoding mapped luma residual values and mapped chroma residual values for said samples; determining a reconstructed luminance value for said sample from said decoded mapped luminance residual value and from said luminance prediction value using both a mapping function and an inverse mapping function that is the inverse of said mapping function; determining reconstructed chrominance values from the mapped chrominance residual values and from the chrominance prediction values using values representing the mapped luminance prediction values; a method comprising: determining the reconstructed luminance values - mapping said luminance prediction value with said mapping function to obtain a mapped luminance prediction value; obtaining an intermediate value representing the sum of said mapped luminance prediction value and said mapped luminance residual value; Including, The method, wherein the reconstructed intensity values are obtained by mapping the intermediate values with the inverse mapping function.
3. 1. An encoding device comprising an electronic circuit, the electronic circuit comprising: - obtaining a luma prediction value of a luma source value of a luma component of a sample of a block of a picture and a chroma prediction value of a chroma source value of a chroma component of said sample, said luma prediction value and said chroma prediction value being obtained by spatial prediction or temporal prediction of said block; determining mapped luminance residual values from the luminance source values of the luminance component and from the luminance prediction values using a mapping function, - mapping said luminance prediction value with said mapping function to obtain a mapped luminance prediction value; - mapping said luminance source values using said mapping function to obtain said mapped luminance source values; determining a mapped luma residual value, the mapped luma residual value being equal to the difference of the mapped luma prediction value from the mapped luma source value; determining mapped chrominance residual values from the chrominance source values of the chrominance components and from the chrominance predicted values using values representing the mapped luma predicted values; encoding and embedding said mapped luma residual values and said mapped chroma residual values into a bitstream; A device configured to:
4. 1. A decoding device comprising an electronic circuit, the electronic circuit comprising: obtaining luminance predictions of the luminance component of the samples of the picture block; obtaining color difference predictions for the color difference components of said samples; - decoding mapped luma residual values and mapped chroma residual values for said samples; determining a reconstructed luminance value for said sample from said decoded mapped luminance residual value and from said luminance prediction value using both a mapping function and an inverse mapping function that is the inverse of said mapping function; determining reconstructed chrominance values from the mapped chrominance residual values and from the chrominance prediction values using values representing the mapped luminance prediction values; an electronic circuit configured to: determining the reconstructed luminance values - mapping said luminance prediction value with said mapping function to obtain a mapped luminance prediction value; obtaining an intermediate value representing the sum of said mapped luminance prediction value and said mapped luminance residual value; Including, The reconstructed luminance values are obtained by mapping the intermediate values with the inverse mapping function.
5. An apparatus comprising the device of claim 3.
6. An apparatus comprising the device of claim 4.
7. 1. A non-transitory information storage medium storing program code instructions implementing a method for encoding, the method comprising: - obtaining a luma prediction value of a luma source value of a luma component of a sample of a block of a picture and a chroma prediction value of a chroma source value of a chroma component of said sample, said luma prediction value and said chroma prediction value being obtained by spatial prediction or temporal prediction of said block; determining mapped luminance residual values from the luminance source values of the luminance component and from the luminance prediction values using a mapping function, - mapping said luminance prediction value with said mapping function to obtain a mapped luminance prediction value; - mapping said luminance source values using said mapping function to obtain said mapped luminance source values; determining a mapped luma residual value, the mapped luma residual value being equal to the difference of the mapped luma prediction value from the mapped luma source value; determining mapped chrominance residual values from the chrominance source values of the chrominance components and from the chrominance predicted values using values representing the mapped luma predicted values; encoding and embedding said mapped luma residual values and said mapped chroma residual values into a bitstream; A non-transitory information storage medium, including:
8. 1. A non-transitory information storage medium storing program code instructions implementing a method for decoding, said method comprising: obtaining luminance predictions of the luminance component of the samples of the picture block; obtaining color difference predictions for the color difference components of said samples; - decoding mapped luma residual values and mapped chroma residual values for said samples; determining a reconstructed luminance value for said sample from said decoded mapped luminance residual value and from said luminance prediction value using both a mapping function and an inverse mapping function that is the inverse of said mapping function; determining reconstructed chrominance values from the mapped chrominance residual values and from the chrominance prediction values using values representing the mapped luminance prediction values; a method comprising: determining the reconstructed luminance values - mapping said luminance prediction value with said mapping function to obtain a mapped luminance prediction value; obtaining an intermediate value representing the sum of said mapped luminance prediction value and said mapped luminance residual value; Including, The non-transitory information storage medium, wherein the reconstructed luminance values are obtained by mapping the intermediate values with the inverse mapping function.
Citation Information
Patent Citations
In-loop block-based image reshaping in high dynamic range video coding
WO2016164235A1