Data encoding method and device, data decoding method and device, image processing device

By predicting image regions and using DCT and CABAC for entropy encoding, the inefficiencies in existing video data encoding and decoding systems are addressed, resulting in improved compression efficiency and quality.

JP7750270B2Active Publication Date: 2025-10-07SONY GROUP CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023132413
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-02-05
Filing Date
2023-08-16
Publication Date
2025-10-07
Estimated Expiration
2039-01-23

AI Technical Summary

Technical Problem

Existing video data encoding and decoding systems face inefficiencies in compressing and decompressing video data due to the limitations of entropy coding techniques, particularly in handling residual image signals and predicting image blocks.

Method used

The proposed solution involves generating residual image signals by predicting image regions and applying Discrete Cosine Transform (DCT) followed by quantization, scanning, and entropy encoding using Context Adaptive Binary Coding (CABAC) to optimize data compression, while ensuring lossless encoding and decoding processes.

Benefits of technology

This approach enhances the efficiency of video data compression by reducing the amount of coded data, improving compression ratios, and maintaining high image quality during decompression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007750270000001
    Figure 0007750270000001
  • Figure 0007750270000002
    Figure 0007750270000002
  • Figure 0007750270000003
    Figure 0007750270000003
Patent Text Reader

Abstract

To provide an encoding method that performs context-adaptive encoding that encodes data representing the magnitude of a data value and a data value code.SOLUTION: In a data encoding method, an ordered array of data values is encoded as data representing the magnitude of a data value and data representing the sign of the data value (2200), the sign of each data value is predicted from properties of one or more other data values in the ordered array for a set of data values that includes at least some of the data values (2210), and a data value code is encoded for the set of data values on the basis of each predicted value (2220).SELECTED DRAWING: Figure 22
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD This disclosure relates to image data encoding and decoding. [Background technology]

[0002] The discussion of the background art provided herein is provided for the purpose of outlining the context in which the present disclosure is made. The work of the named inventors of this application is not expressly or implicitly admitted as prior art to the present disclosure, to the extent that it is described in this background art section, as well as any portion not admitted as prior art at the time of filing.

[0003] Several video data encoding and decoding systems exist that involve transforming the video data into a frequency domain representation, quantizing the frequency domain coefficients, and then applying some form of entropy coding to the quantized coefficients, thereby compressing the video data. To recover a reconstructed version of the original video data, a corresponding decoding or decompression technique is applied. Summary of the Invention

[0004] The present application addresses or alleviates the problems that arise from such processes.

[0005] Aspects and features of the present disclosure are defined in the following claims.

[0006] It should be noted that the above general description and the following detailed description are merely examples of the present invention, and the present invention is not limited to these. [Brief explanation of the drawings]

[0007] The present disclosure, together with further advantages, is best understood by reference to the following detailed description considered in conjunction with the accompanying drawings.

[0008] [Figure 1]1 is a schematic diagram illustrating an audio / video (A / V) data transmission and reception system that uses video data compression and decompression. [Figure 2] FIG. 1 is a schematic diagram illustrating a video display system using decompression of video data. [Figure 3] 1 is a schematic diagram illustrating an audio / video storage system that uses compression and decompression of video data. [Figure 4] FIG. 1 is a schematic diagram illustrating a video camera that uses compression of video data. [Figure 5] FIG. 2 is a schematic diagram illustrating a storage medium. [Figure 6] FIG. 2 is a schematic diagram illustrating a storage medium. [Figure 7] 1 is a schematic diagram illustrating a video data compression and decompression device; [Figure 8] FIG. 2 is a schematic diagram illustrating a prediction unit. [Figure 9] FIG. 1 is a schematic diagram illustrating a partially encoded image. [Figure 10] FIG. 1 is a schematic diagram illustrating a set of possible intra-prediction directions. [Figure 11] FIG. 2 is a schematic diagram illustrating a set of prediction modes. [Figure 12] FIG. 10 is a schematic diagram illustrating another set of prediction modes. [Figure 13] FIG. 1 is a schematic diagram illustrating an intra-prediction process. [Figure 14] FIG. 1 is a schematic diagram illustrating an inter-prediction process. [Figure 15] FIG. 1 is a schematic diagram illustrating an example of a CABAC encoder. [Figure 16] FIG. 2 is a schematic diagram illustrating an array of data values. [Figure 17] FIG. 10 is a schematic diagram illustrating an example of sign bit dependency. [Figure 18] FIG. 1 is a schematic diagram showing an array region. [Figure 19] FIG. 1 is a schematic diagram illustrating a predictor / selector. [Figure 20a] FIG. 1 is a schematic diagram showing a set of relative positions. [Figure 20b] FIG. 1 is a schematic diagram showing a set of relative positions. [Figure 21] FIG. 1 is a schematic diagram illustrating a predictor / selector. [Figure 22] 1 is a schematic flowchart illustrating each method. [Figure 23] 1 is a schematic flowchart illustrating each method. [Figure 24] FIG. 1 is a schematic diagram illustrating a data encoding device. [Figure 25] FIG. 1 is a schematic diagram illustrating a data decoding device. [Figure 26] 1 is a schematic flowchart illustrating each method. [Figure 27] 1 is a schematic flowchart illustrating each method. [Figure 28] FIG. 1 is a schematic diagram showing an example of an array order. DETAILED DESCRIPTION OF THE INVENTION

[0009] Referring now to the drawings, FIGS. 1-4 show schematic diagrams of devices and systems that utilize the compression and / or decompression devices described below in connection with embodiments of the present technology.

[0010] All data compression and / or decompression devices described hereinafter may be implemented as hardware or software running on a general-purpose data processing device such as a general-purpose computer, as programmable hardware such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array), or a combination thereof. In the case of embodiments implemented as software and / or firmware, such software and / or firmware, as well as non-transitory data storage media on which such software and / or firmware is stored or provided, are considered embodiments of the present invention.

[0011] 1 is a schematic diagram illustrating an audio / video data transmission and reception system that uses video data compression and decompression. In this example, the data values ​​to be encoded or decoded represent image data.

[0012] An input audio / video signal 10 is provided to a video data compressor 20 which compresses at least the video component of the audio / video signal 10 for transmission along a transmission route 30, such as a cable, optical fiber, wireless link, etc. The compressed signal is processed by a decompressor 40 to provide an output audio / video signal 50. On the return path, a compressor 60 compresses the audio / video signal for transmission along the transmission route 30 to a decompressor 70.

[0013] Thus, compressor 20 and decompressor 70 may form one node in the transmission link, and decompressor 40 and compressor 60 may form another node in the transmission link. Of course, if the transmission link is unidirectional, only one node needs a compressor, and only the other node needs a decompressor.

[0014] 2 is a schematic diagram illustrating a video display system using video data decompression. In particular, a compressed audio / video signal 100 is processed by a decompressor 110 to provide a decompressed signal that can be displayed on a display device 120. The decompressor 110 may be implemented as an integral part of the display device 120, e.g., located in the same housing as the display device. Alternatively, the decompressor 110 may be provided (for example) as a so-called set-top box (STB). Note that the term "set-top" does not imply that the box is located in a particular orientation or position relative to the display device 120; rather, the term is used simply to indicate a device that can be connected to the display as a peripheral.

[0015] 3 is a schematic diagram illustrating an audio / video storage system using video data compression and decompression. An input audio / video signal 130 is provided to a compressor 140 which produces a compressed signal that is stored by a storage device 150, such as a magnetic disk device, optical disk device, magnetic tape device, or a semiconductor storage device such as a semiconductor memory or other storage device. During playback, the compressed data is read from storage device 150 and passed to a decompressor 160 for decompression, which provides an output audio / video signal 170.

[0016] Of course, a compressed or encoded signal, as well as a storage medium, such as a machine-readable non-transitory storage medium, that stores the signal, is considered an embodiment of the technology.

[0017] Figure 4 is a schematic diagram illustrating a video camera that uses video data compression. In Figure 4, an image capture device 180, such as a CCD (Charge Coupled Device) image sensor and associated control and readout electronics, generates a video signal that is passed to a compressor 190. A microphone (or microphones) 200 generates an audio signal that is passed to the compressor 190. The compressor 190 generates a compressed audio / video signal 210 (represented generally as schematic stage 220) that is stored and / or transmitted.

[0018] The techniques described hereinafter primarily relate to video data compression and decompression. Naturally, many existing techniques may be used for audio data compression in conjunction with the video data compression techniques described hereinafter to generate compressed audio / video signals. Therefore, a separate discussion of audio data compression will not be provided. Also, data rates associated with video data, particularly broadcast-quality video data, are generally much higher than data rates associated with audio data (whether compressed or uncompressed). Therefore, uncompressed audio data may accompany compressed video data to form compressed audio / video signals. Furthermore, although the present example (see FIGS. 1-4) relates to audio / video data, the techniques described hereinafter may also be used in systems that simply handle (i.e., compress, decompress, store, display, and transmit) video data. That is, this embodiment can be applied to video data compression without necessarily processing associated audio data.

[0019] Thus, Figure 4 shows an example of a video capture device including an encoding device of the type described below and an image sensor, while Figure 2 shows an example of a decoding device of the type described below and a display device on which the decoded images are output.

[0020] By combining Figures 2 and 4, a video capture device may be provided that includes an image sensor 180, an encoding device 190, a decoding device 110, and a display device 120 on which the decoded images are output.

[0021] 5 and 6 are schematic diagrams of storage media that store (for example) compressed data generated by apparatus 20, 60, or compressed data input to apparatus 110, storage medium, or stage 150, 220. FIG. 5 is a schematic diagram of a disk storage medium, such as a magnetic disk or optical disk. FIG. 6 is a schematic diagram of a solid-state storage medium, such as flash memory. Note that FIGS. 5 and 6 are also examples of non-transitory machine-readable storage media. The non-transitory machine-readable storage media store computer software that, when executed by a computer, causes the computer to perform one or more of the methods described below.

[0022] That is, the above configurations are examples of video storage, capture, transmission, or reception devices that implement the present technology.

[0023] FIG. 7 is a schematic diagram illustrating a video / image data compression and decompression apparatus for encoding and / or decoding image data representing one or more images.

[0024] The control unit 343 controls the overall operation of the video / image data compression / decompression device. Specifically, the control unit 343 controls the trial encoding process by acting as a selector for selecting various operational aspects, such as block size and shape, with respect to compression modes. The control unit 343 also controls whether or not lossy encoding occurs. This control unit is considered part of the image encoder or image decoder (as the case may be). A series of images of the input video signal 300 are provided to the adder unit 310 and the image predictor unit 320, which will be described in more detail below with reference to FIG. 8. The image encoder or image decoder (as the case may be) may use features of the apparatus shown in FIG. 7 together with the intra-image predictor unit shown in FIG. 8. However, this image encoder or image decoder does not necessarily require all of the features shown in FIG. 7.

[0025] The adder 310 receives the input video signal 300 on its "+" input and the output of the image predictor 320 on its "-" input, effectively performing a subtraction (negative addition) operation, thereby subtracting the predicted image from the input image, resulting in a so-called residual image signal 330 representing the difference between the actual image and the predicted image.

[0026] One reason for generating a residual image signal is as follows: The data coding techniques described, i.e., techniques applied to the residual image signal, tend to work more efficiently when the image to be coded has little "energy." Here, the term "efficient" refers to generating a small amount of coded data. For a particular image quality level, it is desirable (and considered "efficient") to generate as little data as possible. The "energy" in a residual image relates to the amount of information contained in the residual image. If a predicted image and an actual image are identical, the difference between these two images (i.e., the residual image) contains zero information (zero energy) and can be very easily coded into a small amount of coded data. In general, if the prediction process can be performed reasonably well so that the content of the predicted image is similar to the content of the image to be coded, it is expected that the residual image data will contain less information (less energy) than the input image and can be easily coded into a small amount of coded data.

[0027] Thus, encoding (using adder 310) involves predicting an image region of the image to be coded and generating a residual image region that depends on the difference between the predicted image region and the corresponding region of the image to be coded. In conjunction with the techniques described below, an ordered array of data values ​​comprises data values ​​of a representation of the residual image region. Decoding involves predicting an image region of the image to be decoded and generating a residual image region that indicates the difference between the predicted image region and the corresponding region of the image to be decoded. The ordered array of data values ​​comprises data values ​​of a representation of the residual image region. Decoding also includes combining the predicted image region and the residual image region.

[0028] The remaining components of the device operating as an encoder (encoding the residual or difference image) are now described. Residual image data 330 is provided to a transform unit or circuit 340 that generates a Discrete Cosine Transform (DCT) representation of a block or region of residual image data. The DCT technique itself is well known and will not be described in detail here. Furthermore, the use of a DCT is merely illustrative of one possible implementation. Other transforms, such as a Discrete Sine Transform (DST), can be used. Transforms may also be combined, e.g., one transform followed (directly or indirectly) by another. The choice of transform may be explicitly determined or may depend on side information used to configure the encoder and decoder.

[0029] Thus, for example, an encoding and / or decoding method may include predicting an image region of an image to be coded and generating a residual image region that depends on the difference between the predicted image region and a corresponding region of the image to be coded, and an ordered array of data values ​​(described below) containing data values ​​of a representation of the residual image region.

[0030] The output of the transform unit 340, i.e., the set of DCT coefficients for each transform block in the image data, is provided to a quantizer unit 350. A variety of quantization techniques are known in the field of video data compression, ranging from simple multiplication by a quantization scaling factor to the application of complex lookup tables under the control of a quantization parameter. They generally serve two purposes: first, to reduce the number of possible values ​​of the transform data through the quantization process; and second, to increase the likelihood that the transform data will have a zero value through the quantization process. This allows the entropy coding process, described below, to be more efficient in generating small amounts of compressed video data.

[0031] A data scanning process is applied by the scanning unit 360. The purpose of the scanning process is to reorder the quantized transformed data in order to group together as many non-zero quantized transform coefficients as possible, and, of course, to group together as many zero-valued coefficients as possible. These functions allow the efficient application of so-called run-length coding or similar techniques. The scanning process therefore involves selecting coefficients from the quantized transformed data, and in particular, from blocks of coefficients corresponding to blocks of transformed and quantized image data, in a "scan order" such that (a) all coefficients are selected once as part of the scan, and (b) the scan achieves the desired reordering. One example of a scan order that produces useful results is the so-called Up-right Diagonal scan order.

[0032] The scanned coefficients are then passed to an entropy encoder (EE) 370. Again, various entropy encoding techniques may be used. Two examples are variations of the so-called Context Adaptive Binary Coding (CABAC) system and variations of the so-called Context Adaptive Variable-Length Coding (CAVLC) system. CABAC is generally considered efficient. Studies have shown that the amount of coded output data in CABAC is 10-20% less than that of CAVLC for comparable image quality. However, the implementation complexity of CAVLC is considered much lower than that of CABAC. Note that although the scanning and entropy encoding processes are shown as separate processes, they may actually be combined or handled together. That is, data may be read into the entropy encoder in scan order. This also applies to the inverse processes described below.

[0033] The output of entropy encoder 370 provides compressed output video signal 380, together with additional data (described above and / or below) that specifies, for example, how predictor 320 generated the predicted image.

[0034] However, since the operation of the predictor 320 itself depends on the decompressed compressed output data, a return path is also provided.

[0035] The reason for this functionality is as follows: At an appropriate stage in the decompression process (described below), decompressed residual data is generated. This decompressed residual data must be added to a predicted image to generate the output image (because the original residual data is the difference between the input image and the predicted image). To make this process equivalent on the compression and decompression sides, the predicted image generated by the predictor 320 should be the same during the compression and decompression processes. Of course, the device does not have access to the original input image during decompression; it only has access to the decompressed image. Therefore, during compression, the predictor 320 bases its predictions (at least for inter-image coding) on ​​the decompressed compressed image.

[0036] The entropy encoding process performed by the entropy encoder 370 may be considered (at least in some instances) to be "lossless," i.e., to return exactly the same data that was originally provided to the entropy encoder 370. Therefore, in such instances, a return path may be implemented prior to the entropy encoding stage. Indeed, the scanning process performed by the scanner 360 may also be considered lossless, although in this embodiment, the return path 390 runs from the output of the quantizer 350 to the input of the supplemental inverse quantizer 420. If a stage is or may be lossy, that stage may be included in the feedback loop formed by the return path. For example, the entropy encoding stage may be lossy, at least in principle, by, for example, a technique for encoding bits in parity information. In such instances, the entropy encoding and decoding must be included as part of the feedback loop.

[0037] In general, entropy decoder 410, inverse scan unit 400, inverse quantizer unit 420, and inverse transform unit or circuit 430 provide the inverse functions of entropy encoder 370, scan unit 360, quantizer unit 350, and transform unit 340, respectively. The compression process will be described here; the process for decompressing the input compressed video signal will be described separately below.

[0038] During the compression process, the scanned coefficients are passed from the quantizer 350 via a return path 390 to an inverse quantizer 420, which performs the inverse operation of the scanner 360. The inverse quantization and inverse transform processes are performed by the inverse quantizer 420 and the inverse transformer 430, respectively, to produce a compressed-decompressed residual image signal 440.

[0039] Image signal 440 is summed with the output of prediction unit 320 in summing unit 450 to produce a reconstructed output image 460, which forms one of the inputs to image prediction unit 320, as described below.

[0040] The process applied to decompress a received compressed video signal 470 will now be described. The compressed video signal 470 is first fed to an entropy decoder 410, which in turn feeds an inverse scan unit 400, an inverse quantization unit 420, and an inverse transform unit 430. It is then summed with the output of the image prediction unit 320 by an adder 450. Thus, at the decoder side, the decoder reconstructs a residual image and applies it (block by block) to the predicted image (by the adder 450) to decode each block. In short, the output 460 of the adder 450 forms the output decompressed video signal 480. In practice, further filtering (e.g., using a filter 560) may optionally be performed before the signal is output. This filter 560 is shown in FIG. 8. In comparison with FIG. 8, the filter 560 has been omitted from the overall configuration of FIG. 7 for clarity.

[0041] The devices shown in Figures 7 and 8 can operate as either compression (encoding) devices or decompression (decoding) devices. The functionality of the two types of devices substantially overlaps. Scanner 360 and Entropy Encoder 370 are not used in decompression mode. Predictor 320 (described in more detail below) and other components operate according to mode and parameter information contained in the received compressed bitstream, rather than generating this information themselves.

[0042] FIG. 8 is a schematic diagram illustrating the generation of a predicted image, and in particular the operation of the image predictor 320.

[0043] Two basic prediction modes are performed by the image prediction unit 320: the so-called intra-image prediction and the so-called inter-image prediction or motion-compensated (MC) prediction. On the encoder side, these predictions each involve finding a prediction direction for the current block to be predicted and generating a predictive block of samples depending on other samples (in the same (intra) or another (inter) image). The adder 310 or 450 encodes or decodes, respectively, the block by encoding or decoding the difference between the predicted block and the actual block.

[0044] (At the decoder side or inverse decoding side of the encoder, this detection of the prediction direction may depend on data associated with the data encoded by the encoder that indicates which direction was used by the encoder, or the detection of the prediction direction may depend on the same factors determined by the encoder.)

[0045] Intra-image prediction is based on predicting the content of blocks or regions of an image on data obtained from within the same image. This corresponds to the so-called I-frame coding in other video compression techniques. However, in contrast to I-frame coding, in which the entire image is coded by intra-coding, in this embodiment the choice between intra-coding and inter-coding can be made on a block-by-block basis. In other embodiments, the choice is made on a picture-by-picture basis.

[0046] Motion compensated prediction is an example of inter-image prediction, in which motion information is used that attempts to define the source, in other adjacent or nearby images, of the image detail to be coded in the current image. Thus, in an ideal case, the content of a block of image data in a predicted image can very easily be coded as a reference (motion vector) that points to a corresponding block in the adjacent image at the same or a slightly different position.

[0047] A technique known as "block copy" prediction is in some ways a hybrid of the two above, as it uses a vector that indicates a block of samples displaced from the current predicted block within the same image that should be copied to generate the current predicted block.

[0048] Returning to FIG. 8, two image prediction configurations (corresponding to intra-image prediction and inter-image prediction) are shown, which are selected by multiplier unit 500 under control of mode signal 510 (e.g., from control unit 343) to provide a block of predicted images for supply to adders 310 and 450. The selection is based on which selection results in the least "energy" (which, as discussed above, can be thought of as the amount of information that needs to be coded), and this selection is communicated to the decoder in the coded output data stream. In this regard, image energy can be detected, for example, by performing trial subtractions of areas of two versions of the predicted image from the input image, squaring each pixel value of the difference image, summing the squared values, and identifying which of the two versions has a lower mean-squared difference image value associated with that image area. In other examples, trial coding can be performed for each selection or for each possible selection. The selection is then made according to the cost of each possible selection in terms of the number of bits required for coding and / or the distortion to the image.

[0049] In an intra-coding system, the actual prediction is based on the image blocks received as part of signal 460. That is, the prediction is based on the encoding-decoding image blocks so that the exact same prediction can be made in the decompressor. However, data can also be derived from the input video signal 300 by the intra mode selector 520 to control the operation of the intra image predictor 530.

[0050] In inter-image prediction, a motion compensated (MC) predictor 540 uses motion information, such as motion vectors, derived from the input video signal 300 by a motion estimation unit 550. The MC predictor 540 applies these motion vectors to the processed reconstructed image 460 to generate blocks of inter-image prediction.

[0051] Thus, the intra-image prediction unit 530 and the motion compensation prediction unit 540 (operating together with the estimation unit 550) respectively operate as a detection unit that detects the prediction direction for the current block to be predicted, and as a generation unit that generates a predicted block of samples (which form part of the prediction result passed to the addition units 310 and 450) depending on other samples defined by the prediction direction.

[0052] The processing applied to signal 460 will now be described. First, signal 460 is optionally filtered by filter unit 560, which will be described in more detail below. This processing involves applying a "deblocking" filter to eliminate or at least reduce the impact on the block-based processing and subsequent operations performed by transform unit 340. A sample adaptive offset (SAO) filter may also be used. An adaptive loop filter is also optionally applied, using coefficients obtained by processing reconstructed signal 460 and input video signal 300. This adaptive loop filter is a type of filter that applies adaptive filter coefficients to the data to be filtered, using known techniques. That is, the filter coefficients may vary based on various factors. Data defining which filter coefficients to use is included as part of the encoded output data stream.

[0053] When the device is operating as a decompressor, the filtered output from filter 560 actually forms output video signal 480. This signal is buffered in one or more image or frame stores 570. Storage of a series of images is required for the motion compensation prediction process, particularly for the generation of motion vectors. To conserve storage space, the images stored in image store 570 may be held in a compressed format and then decompressed for use in generating motion vectors. Any known compression / decompression system may be used for this specific purpose. The stored images are passed to interpolation filter 580, which generates higher resolution stored images. In this example, intermediate samples (sub-samples) are generated so that the interpolated image output by interpolation filter 580 has four times the resolution (in each dimension) of the image stored in image store 570 if the luminance channel is 4:2:0, and eight times the resolution (in each dimension) of the image stored in image store 570 if the chrominance channels are 4:2:0. The interpolated images are passed as input to motion estimation unit 550 and motion compensation prediction unit 540.

[0054] We now describe how to divide an image for compression purposes. At a basic level, the image to be compressed can be thought of as an array of blocks or regions of samples. This division of an image into blocks or regions can be achieved by decision trees, as described in the following documents, the contents of which are incorporated herein by reference: SERIES H: AUDIOVISUAL AND MULTIMEDIA SYSTEMS Infrastructure of audiovisual services - Coding of moving video High efficiency video coding Recommendation ITU-T H.265 12 / 2016. Also, High Efficiency Video Coding (HECV) algorithms and Architectures (authors Madhukar Budagavi, Gary J. Sullivan, Vivienne Sze; ISBN 978-3-319-06894-7; 2014) is incorporated herein by reference. In some examples, the resulting blocks or regions have varying sizes and, in some cases, shapes that generally follow the alignment of image features within the image, as determined by a decision tree. This alone can improve coding efficiency, since samples that represent or are aligned with similar image features tend to be grouped together in such an arrangement. In some examples, square blocks or regions of different sizes (e.g., 4x4 samples to, e.g., 64x64, or larger blocks) are available for selection. In other implementations, blocks or regions of different shapes may be used, such as rectangular blocks (e.g., vertically or horizontally oriented). Other non-square and non-rectangular blocks may also be used. As a result of such division of the image into blocks or regions, (at least in this example) each sample of the image is assigned to one, and even only one, such block or region sample.

[0055] The intra prediction process will now be described. Generally, intra prediction involves generating a prediction of a current block of samples from previously coded and decoded samples in the same image.

[0056] Figure 9 is a schematic diagram showing a partially coded image 800, where the image is coded block by block from the top left to the bottom right. An example of a block that is partway through coding in the processing of the entire image is shown as block 810. The upper shaded region 820 to the left of block 810 has already been coded. Any of the shaded regions 820 can be used for intra-image prediction of the content of block 810, but the unshaded regions below cannot.

[0057] In some examples, an image is coded block-by-block, such that larger blocks (called CUs (coding units)) are coded in an order such as that described with reference to FIG. 9 . For each CU, the CU may be treated as a collection of two or more smaller blocks or transform units (TUs) (depending on the block partitioning process performed). This provides a hierarchical coding order in which the image is coded in units of CUs. Each CU may be coded in units of TUs. However, for each TU (the largest node in the block partitioning tree structure) within the current coding tree unit, the hierarchical coding order described above (CU-by-CU, then TU-by-TU) means that previously coded samples in the current CU may exist and be available for coding the TU. These may be located, for example, to the top right or bottom left of the TU.

[0058] Block 810 represents a CU. As mentioned above, this may be subdivided into a collection of smaller units for intra-image prediction processing. An example of a current TU 830 is shown in CU 810. More generally, an image is divided into regions or groups of samples so that signaling information and transformed data can be efficiently coded. Signaling information may require a different tree structure consisting of subdivisions for transform information, and indeed, prediction information or prediction itself. For this reason, a coding unit may have a different tree structure for transform blocks or regions, prediction blocks or regions, and prediction information structures. In some examples, the structure may be a so-called quadtree coding unit, where leaf nodes contain one or more prediction units and one or more transform units. A transform unit may contain multiple transform blocks corresponding to the luminance and color representations of an image, and prediction may be considered to be applicable at the transform block level. In some examples, parameters applied to a particular group of samples may be considered to be defined primarily at the block level, which may not have the same granularity as the transform structure.

[0059] Intra-picture prediction considers samples coded before considering the current TU. These samples are samples above and / or to the left of the current TU. The source samples from which the required samples are predicted may be in different positions or orientations relative to the current TU. To determine which direction is suitable for the current prediction unit, the mode selector 520 of an exemplary encoder may try all available TU structure combinations for each candidate direction and select the prediction direction and TU structure that provides the best compression efficiency.

[0060] An image may be coded in "slices." In one example, a slice is a group of horizontally adjacent CUs. However, more generally, a slice may comprise the entire residual image, or it may be a single CU or a row of CUs, etc. Because slices are coded as independent units, they provide a degree of error resilience. The encoder and decoder states are completely reset at slice boundaries. For example, intra prediction is not performed across slice boundaries. Therefore, slice boundaries are treated as image boundaries.

[0061] Figure 10 is a schematic diagram showing a set of possible (candidate) prediction directions. A set of all candidate directions is available to the prediction unit. The direction is determined by moving horizontally and vertically relative to the current block position and is coded as a prediction "mode". This set of prediction modes is shown in Figure 11. Note that the so-called DC mode represents a simple arithmetic mean of the surrounding top and left samples. Also, the set of directions shown in Figure 10 is only an example. Another example is a set of (for example) 65 angular modes plus DC and planar (for a total of 67 modes), as shown schematically in Figure 12. Other numbers of modes are also possible.

[0062] In general, the system is operable, after detecting the prediction direction, to generate a predictive block of samples depending on other samples defined by the prediction direction. In some examples, the image encoder is configured to encode data identifying the selected prediction direction for each sample or region of the image (and the image decoder is configured to detect such data).

[0063] Figure 13 is a schematic diagram illustrating an intra-prediction process, in which a sample 900 of a block or region 910 of samples is derived from other reference samples 920 of the same image according to a direction 930 defined by the intra-prediction mode associated with the sample. In this example, the reference samples 920 originate from blocks above and to the left of the current block 910, and a predicted value for the sample 900 is obtained by tracking the reference samples 920 along the direction 930. The direction 930 may indicate a single individual reference sample, but more commonly, an interpolated value of surrounding reference samples is used as the predicted value. Note that the block 910 may be square, as shown in Figure 13, or may have other shapes, such as a rectangle.

[0064] Figure 14 is a schematic diagram illustrating an inter-prediction process in which a block or region 1400 of a current image 1410 is predicted with respect to blocks 1430, 1440, or both, pointed to by a motion vector 1420 in one or more other images 1450, 1460. The blocks predicted in this way can be used to generate residual data that is coded as described above.

[0065] CABAC encoding FIG. 15 is a schematic diagram illustrating the operation of the CABAC entropy encoder.

[0066] In some embodiments, context-adaptive coding of this function may involve encoding small amounts of data with respect to a probability model, i.e., a context, that represents the expected or predicted likelihood of a data bit being 1 or 0. To this end, input data bits are assigned a code value within a selected one of two (or, more generally, multiple) complementary subranges of various code values, the size of each subrange (and, in some embodiments, the proportion of each subrange to the set of code values) being defined by a context (which is in turn associated with the input value or defined by a context variable for the input value). In a next step, the total range, i.e., the set of code values ​​(to be used for the next input data bit or value), is modified according to the assigned code value and current size of the selected subrange. If the modified range is then smaller than a threshold representing a predetermined minimum size (e.g., half the original range size), it is increased in size, e.g., by doubling (shifting left) the modified range. This doubling process can be performed successively (two or more times), if necessary, until the range reaches at least the predetermined minimum size. At this point, an output encoded data bit is generated, thereby indicating that a doubling or size increase operation(s) has occurred. In another step, the context is modified (i.e., in embodiments, the context variable is modified) for use with or for the next input data bit or value (in some embodiments, for the next group of data bits or values ​​to be encoded). This may be done using the current context and an identifier of the current "most likely symbol" (1 or 0, whichever the context indicates currently has a probability greater than 0.5) as an index into a lookup table of new context values ​​or as input to an appropriate mathematical formula to derive a new context variable. In embodiments, modifying the context variable increases the proportion of the set of code values ​​in the subrange selected for the current data value.

[0067] CABAC encoders operate on binary data, i.e., data represented by only two symbols: 0 and 1. They use a so-called context modeling process that selects a "context," i.e., a probability model, for the upcoming data based on data already encoded. The context selection is performed in a deterministic way, based on data already decoded, so that the same decision can be made at the decoder without requiring additional data (that identifies the context) to be added to the encoded data stream passed to the decoder.

[0068] Referring to Figure 15, input data to be encoded may be passed to a binary converter 900 unless it is already in binary form. If the data is already in binary form, converter 900 is bypassed (by switch 910 in the figure). In this embodiment, the conversion to binary form is actually accomplished by representing the quantized transform coefficient data as a series of binary "maps," which are described below.

[0069] The binary data may then be processed by one of two processing paths: a "normal" path and a "bypass" path (shown schematically as separate paths, but which in the embodiments described below may actually be implemented in the same processing stage with only slightly different parameters). The bypass path employs a so-called bypass coder 920 that does not necessarily utilize the same form of context modeling as the normal path. In some instances of CABAC encoding, the bypass path may be selected when a batch of data needs to be processed particularly quickly. However, in this embodiment, two characteristics of the so-called "bypass" data are noted. First, bypass data is processed by the CABAC encoder (950, 960) utilizing only a fixed context model representing a 50% probability. Second, bypass data relates to certain categories of data. A particular example of such data is coefficient code data. If the bypass path is not selected, the normal path is selected by the illustrated switches 930, 940. This includes data processed by the context modeler 950 and subsequently by the coding engine 960.

[0070] The entropy encoder shown in FIG. 15 encodes a block of data (i.e., data corresponding to a block of coefficients for a block of a residual image) as a single value if the entire block consists of zero-valued data. For each block that does not fall into this category, i.e., a block that contains at least some nonzero data, a "significance map" is created. The significance map indicates, for each position in the block of data to be encoded, whether the corresponding coefficient in the block is nonzero. (Thus, this is an example of a significance map that indicates the position of the most significant data portion that is nonzero for a data value array.) The significance map may also include a data flag indicating the position of the last most significant data portion having a nonzero value in a predetermined order in the data value array. The order of the data value array for this encoding may be, for example, from smallest to largest spatial frequency coefficients. This may differ from the array order for encoding the data value codes, which may, for example, be from the last nonzero coefficient (in increasing spatial frequency order) to the largest spatial frequency coefficient.

[0071] The significance map data itself, which is in binary form, is CABAC coded. Using the significance map aids in compression because data does not need to be coded for coefficients whose magnitudes are indicated as zero by the significance map. The significance map can also contain a special code to indicate the last non-zero coefficient in a block. This allows all of the last high-frequency / trailing-zero coefficients to be omitted from coding. In the coded bitstream, the significance map is followed by data defining the values ​​of the non-zero coefficients specified by the significance map.

[0072] Additional levels of map data are also created for CABAC encoding. One example is a map defining, as a binary value (1=yes, 0=no), whether the coefficient data at map locations marked as "non-zero" by the significance map actually have a value of "1." Another map identifies whether the coefficient data at map locations marked as "non-zero" by the significance map actually have a value of "2." Yet another map indicates, for those map locations marked as "non-zero" by the significance map, whether the data has a value of "3 or greater." Yet another map indicates, for data identified as "non-zero," the sign of the data value (using a predetermined binary notation, such as 1 for +, 0 for -, or vice versa).

[0073] In some embodiments, the significance map and other maps are generated from quantized transform coefficients, for example by the scanning unit 360, and subjected to a zigzag scan process (or a scan process selected from zigzag scan, horizontal raster scan, and vertical raster scan for intra prediction modes), followed by CABAC encoding.

[0074] In some embodiments, the CABAC entropy coder encodes syntax elements by the following process.

[0075] The location of the last significant coefficient within a TU (in scan order) is coded.

[0076] For each 4x4 coefficient group (these groups are processed in reverse scan order), we code a significant coefficient group flag that indicates whether the group contains a non-zero coefficient. This is not necessary for the last group containing significant coefficients, and is assumed to be 1 for the top-left group (which contains the DC coefficient). If the flag is 1, it is immediately followed by coding the following syntax elements for that group:

[0077] Efficacy Map: For each coefficient in the group, a flag is coded to indicate whether the coefficient is valid (has a non-zero value). No flag is needed for the coefficient indicated in the last valid position.

[0078] 2 or more maps: For up to eight coefficients (counting backwards from the end of the group) with a value of 1 in the significance map, indicate whether their magnitude is 2 or greater.

[0079] 3 or more flags: For one coefficient with a value of 1 in the 2 or greater map (closest to the end of the group), indicates whether its magnitude is 3 or greater.

[0080] Sign bit: In the above configuration, the code bits are coded as equal probability CABAC bins. The last code bit (in reverse scan order) is inferred from parity when using hidden code bits, possibly. In addition to the above configuration, configurations applicable to embodiments of the present disclosure are described below.

[0081] Escape Codes: For coefficients whose magnitude cannot be fully represented by earlier syntax elements, the remainder is encoded as an escape code.

[0082] Generally, CABAC coding involves predicting a context, i.e., a probability model, for the next bit to be coded based on other data that has already been coded. If the next bit is the same as the bit identified as "most probable" by the probability model, coding the information "the next bit matches the probability model" can be done very efficiently. In comparison, coding the information "the next bit does not match the probability model" is inefficient. Therefore, deriving context data is important for successful operation of the encoder. The term "adaptive" means that the context, i.e., the probability model, is applied or changed during coding in an attempt to better adapt to the next data (which has not yet been coded).

[0083] As a simple example, in written English, the letter "U" is relatively rare. However, immediately following the letter "Q," "U" does occur frequently. Thus, a probabilistic model might set the probability of "U" to a very low value, but if the current letter is "Q," the probabilistic model for "U" as the next letter could set a very high probability.

[0084] In this configuration, CABAC coding is used for at least the significance map and the map indicating whether a non-zero value is 1 or 2. However, each of these syntax elements does not necessarily have to be coded for each coefficient. In these embodiments, bypassing is the same as CABAC coding, except that the probability model is fixed at a probability distribution where 1 and 0 are equal (0.5:0.5), and is used for at least the code data and some portion of the coefficient magnitude not described by the preceding syntax element. For those data locations identified as having some portion of the coefficient magnitude not fully described, so-called escape data coding can be used individually to code the actual remaining value of the data, where the actual magnitude value is this remaining magnitude value plus an offset derived from each coded syntax element. This may include Golomb-Rice coding techniques.

[0085] The CABAC context modeling and encoding process is described in detail in SERIES H: AUDIOVISUAL AND MULTIMEDIA SYSTEMS Infrastructure of audiovisual services - Coding of moving video High efficiency video coding Recommendation ITU-T H.265 12 / 2016, and in High Efficiency Video Coding (HECV) algorithms and Architectures (authors Madhukar Budagavi, Gary J. Sullivan, Vivienne Sze; ISBN 978-3-319-06894-7; 2014, Chap 8 pp. 209-274), which is also incorporated herein by reference.

[0086] Sign bit encoding In the embodiments described below, the code bits may be coded using other than equiprobable coding, such as predicting the code of each data value for a set of data values ​​that includes at least some of the data values ​​from properties of one or more other data values ​​in the ordered array, and coding the code of the data values ​​for the set of data values ​​based on the predicted values.

[0087] To achieve this, a predictor of at least one sign bit is used. Note that in obtaining this predictor, the data values ​​partially represented by the sign bits comprise apparently independent data values ​​(e.g., spatial frequency coefficients), but the data value signs may be at least partially dependent. The origins of this dependency and empirical observations are discussed below. The origins of this dependency and empirical observations influence the design and operation of embodiments of the present disclosure.

[0088] It should be noted that accurate prediction of whether a particular code bit will be a minus or plus sign bit is not actually necessary to improve coding efficiency. Indeed, when using, for example, a context-adaptive encoder or similar type of arrangement that takes into account the probability of the next symbol to be encoded, all that is needed is the ability to predict whether a particular code bit is more than 50% likely to be a minus sign bit or more than 50% likely to be a plus sign bit. Using such a prediction in a context-adaptive encoder, such as the one shown in Figure 15, allows for the selection of an appropriate context for encoding the next code bit.

[0089] We now turn to some background information on sign bit prediction: First, we will discuss examples related to DCT coding, but in the following the technique will be extended to other configurations such as the discrete sine transform (DST), the so-called transform skip mode, and the so-called non-separable quadratic transform (NSST) mode.

[0090] The non-separable quadratic transform (NSST) is described in "Algorithm Description of Joint Exploration Test Model 1," Chen et al., Joint Video Exploration Team (JVET) document JVET-A1001, which discloses the NSST representing a quadratic transform applied between the forward core transform and quantization (at the encoder side) and between dequantization and the inverse core transform (at the decoder side). The contents of this document are incorporated herein by reference.

[0091] Thus, in some implementations, the ordered array of data values ​​includes data values ​​of a frequency transform representation of the residual image region. The frequency transform may include, for example, a discrete cosine transform (DCT), a discrete sine transform (DST), a DCT in one direction and a DST in an orthogonal direction, and a linear transform followed by a non-separable quadratic transform. In other implementations, the ordered array of data values ​​may include data values ​​of a reordered representation of the residual image region in transform skip mode.

[0092] The transform skip operation is described in section 8.6.2 of The CABAC context modeling, and details of the coding process are described in SERIES H: AUDIOVISUAL AND MULTIMEDIA SYSTEMS Infrastructure of audiovisual services - Coding of moving video High efficiency video coding Recommendation ITU-T H.265 12 / 2016. Also, High Efficiency Video Coding (HECV) algorithms and Architectures (by Madhukar Budagavi, Gary J. Sullivan, and Vivienne Sze; ISBN 978-3-319-06894-7; 2014, Chap 8, pp. 209-274) is incorporated herein by reference. For some regions or blocks, coding gain can be achieved by skipping the transform. The residual in the spatial domain is quantized and coded. In some cases, the spatial order may be reversed (which allows for an adjustment of the predicted magnitude of the spatial difference, since the predicted order of coefficient magnitudes in the frequency transform block is such that the largest magnitude is in the bottom right block furthest from the reference sample (typically the largest magnitude is in the top left corner of the DC coefficient).

[0093] For intra-predicted data in a residual block, it is generally expected that the closer the residual sample is to each reference sample, the smaller the error or magnitude of the residual value. (In contrast, for inter-predicted data, the error is more likely to be uniform across the block.) According to the notation used here and frequently used in other descriptions of this technology, residual samples closer to the top left of a residual block tend to have less error than residual blocks at the bottom and / or right of the residual block. Indeed, it is partly for this reason that we have proposed a so-called "short distance intra-prediction (SDIP)" configuration.

[0094] Short-distance intra prediction (SDIP) is described in CE6.b1 Report on Short Distance Intra Prediction Method, Cao et al., Joint Collaborative Team on Video Coding (JCT-VC) document: JCTVC-E278. In SDIP, large blocks are divided into so-called SDIP partitions, and each SDIP partition can be coded separately. The contents of this document are incorporated herein by reference.

[0095] However, considering the frequency domain extrapolation for the DCT example, the schematic diagram in Figure 16 shows the set of DCT basis functions for an example 8x8 block, where the DC coefficients have increasing horizontal spatial frequency moving from the top left to the right, and increasing vertical spatial frequency moving downward.

[0096] Taking the horizontal direction as an example, the energy in the residual block decreases towards the left of the block, so that DCT coefficient 1600 (coefficient (0,0)) is likely to have a different sign than its horizontally adjacent coefficient (0,1) 1610. That is, sign(c 0,0 ) is the sign (c 0,1 ) is less than 0.5.

[0097] This is just one example for one set of coefficient locations. In fact, empirical studies using large sets of test data have shown that correlations (i.e., positive or negative correlations) can exist between code bits across sets of locations within a DCT block.

[0098] Inter-coding is described below. When intra-coding is used to generate a prediction value for obtaining a residual block, the characteristics of this correlation may depend on the prediction mode or direction used. Above, examples of sets of prediction modes have been described with reference to Figures 10 to 12. Although this principle is applicable to any set, the following description uses the example set shown in Figure 11. Thus, in various examples, predicting an image region of an image to be coded includes predicting a sample of the image region based on other coded and decoded samples of the image that are shifted from the prediction sample in a direction specified in the prediction mode.

[0099] Consider the example of a common vertical prediction mode (e.g., mode 25, margins such as 25+ / -4), where dependencies between code bits tend to generally (but not necessarily) act vertically. Referring to Figure 17, we use the example of one location within a block (location 1700, denoted by X), but of course similar principles are (at least potentially) applicable to all locations within the block. The code bit that is considered to be most highly correlated with the code of location X in vertical mode is denoted as location 1710, denoted by a "V."

[0100] Consider now the example of a common horizontal prediction mode (e.g., mode 10, margins such as 10+ / -4), where dependencies between code bits generally (but not necessarily) tend to act horizontally. Referring to Figure 17, the code bit that is considered to be most highly correlated with the code at location X in horizontal mode is shown as location 1720, marked with an "H" symbol.

[0101] Consider now the example of a common diagonal prediction mode (e.g., mode 18, 18+ / -4 margin, etc., or mode 2+margin, or mode 34-margin), where the dependency between code bits tends to generally (but not necessarily) act diagonally or in a mixed vertical and horizontal direction. Referring again to Figure 17, the code bit that is considered to be most highly correlated with the code at location X in the diagonal mode is shown as location 1730, marked with a "D" symbol.

[0102] Consider now the example of a DC or planar prediction mode (e.g., mode 0 or 1). Referring again to Figure 17, the sign bit that is considered to be most highly correlated with the sign of location X in DC or planar mode is shown as location 1740, marked with a "P".

[0103] Thus, for each data value in the set of data values, the relative position of one or more other data values ​​from which the data value is predicted depends on the prediction mode applicable to the array of data values.

[0104] When encoding a block using this technique, it may be useful to generate a prediction for a particular code bit depending on previously encoded code bits, allowing for complementary processing in the encoder and decoder, where the corresponding prediction can be generated at the decoder based on the code bits already decoded.

[0105] This technique is applicable to the DCT-coded intra image examples described herein, as well as DST, transform skip, and inter-coded images. Lookup may vary between these different configurations, but the basic technique is equally applicable. A non-separable quadratic transform (NSST) or an extended multi-layer transform (EMT) may be used. Where one transform (e.g., DCT or DST) can be used in one axis (e.g., horizontal or vertical), another transform (e.g., DCT or DST) can be used in the other orthogonal direction.

[0106] For inter-coded blocks, different correlations may be observed, as described below. In the example block shown in Figure 16, the ordered array of data values ​​includes data values ​​of a frequency-transformed representation of a residual image region, which is a region that has already undergone a series of one or more frequency transforms (e.g., a transform and an NSST). For example, the frequency transforms may include one or more of a discrete cosine transform (DCT), a discrete sine transform (DST), a DCT in one direction and a DST in an orthogonal direction, and a linear transform followed by a non-separable quadratic transform. Alternatively, the ordered array of data values ​​may include data values ​​of a reordered representation of the residual image region in transform skip mode.

[0107] Examples of how these correlations can be used are described below.

[0108] Referring to Figure 18, the code prediction process may be performed using different sets of coefficients and different algorithms for different regions of the block to be coded. Area 0 is the upper left (DC) coefficient, Area 1 is the top row, Area 2 is the left column, Area 3 represents all other locations.

[0109] Such grouping may affect the relative locations of the samples and the samples for which the signs are predicted. For each data value in the set of data values, the relative locations of the one or more other data values ​​from which the data value is predicted depends on the location in the array of data values ​​of the data value being predicted. That is, in some examples, for each data value in the set of data values, the relative locations of the one or more other data values ​​from which the data value is predicted depends on the location in the array of data values ​​of the data value being predicted.

[0110] Figure 19 is a schematic diagram of a predictor / selector unit 1900 that receives as input one or more of the following: Currently applicable prediction modes Location of coefficients in the current block Indication of whether the current block is inter-coded or intra-predicted Indication of whether the current block is DCT coded, DST coded, DCT coded in one direction, DST coded in the orthogonal direction, or transform-skip coded A code represented by other coded bits Examples include:

[0111] These inputs are also applied to a lookup table. For example, as a result, the predictor / selector unit 1900 may produce one or both of the following as output 1910: the predicted value of the current sign bit, and / or The context used to encode (or decode in the case of a decoder) the current code bit

[0112] Note that the lookup may refer to other data. In some examples, for each data value in the set of data values, one or more properties of other data values ​​that predict the data value include: Prediction modes applicable to data value arrays, The array size of the data value array, The shape of the data value array, and The location of the data value within the data value array The parameter depends on one or more elements selected from the list consisting of:

[0113] It should be noted that embodiments are not limited to configurations in which the properties of the one or more other data values ​​(based on which predictions or context selections are made) include the signs of the one or more other data values. In other examples, the properties of the one or more other data values ​​may instead or additionally include the magnitudes of the one or more other data values.

[0114] Figures 20a and 20b are schematic diagrams illustrating another example, in which for any particular location X in a block, such as location 2000, the code applicable to this location is predicted (or a context for encoding the code bit is generated) as a lookup function of the codes of other disjoint sets of coded (or decoded in the case of a decoder) code bits A to F at each relative location, which actually exist in the coded or decoded data for location X.

[0115] For each data value in the set of data values, the one or more other data values ​​that predict the data value are located at a predetermined relative position to the data value in the array of data values. Figure 20b shows an example of the above.

[0116] For inter-coded blocks, different correlations may be observed. For inter prediction of DCTxDCT code bits, It is the neighborhood of the current coefficient X (and position x, y). Consider the pattern containing the data value X to be predicted in FIG. 20b. (A) XBC DE F For transform skip blocks (which may be intra- or inter-coded), If (B<0 and D<0) or (B>0 and D>0) Predicted sign = same as B (or D). Else if (B!=0 and D==0) Predicted sign = same as B. Else if (B==0 and D!=0) Predicted sign = same as D. Else if (B!=0 and D!=0 and E!=0) Predicted sign = opposite of E. Else Don't predicted sign - use equiprobable bin. For the inter-conversion block, Let Right=(DCT horizontally) ? C : 0 Below=(DCT vertically) ? F : 0 (That is, when the DST transform is applied in a particular direction, it assumes no correlation with the sign bit and treats the coefficient as if it does not exist.) if (Right<0 and Below<0) || (Right>0 and Below>0) (The symbol || represents the logical OR operation) Predicted sign = opposite sign to Right (or Below - they're the same) else if (Right!=0 and Below ==0) (the != symbol means not equal) Predicted sign = opposite sign to Right else if (Below!=0 and Right ==0) Predicted sign = opposite sign to Below else if (Right!=0 and Below!=0) if (abs(Right)>abs(Below)) Predicted sign = opposite sign to Right. Else Predicted sign = opposite sign to Below. else If (y==0 and (x AND 1)!=0 && (x*2>=width) and DCT horizontally) (&& is a logical operation that returns a boolean value of true if both operands are true, and false otherwise. Also, the logical AND operation finds the location in an 8x8 block where x=5 or 7 and y=0.) Predicted sign = negative sign. Else if (x==0 and (y AND 1)!=0 && (y*2>=height) and DCT vertically) Predicted sign = negative sign. Else Don't predict sign: use equiprobable bin instead.

[0117] In at least some embodiments, the various predictions made above may have different probability models associated with them (i.e., using different contexts) depending on the match.

[0118] For example, if Right and Below have the same sign, it is more likely to be a good prediction than if just Right or Below is present. Therefore, different CABAC contexts can be used in different situations. In some embodiments, which CABAC context to use may depend on the location (e.g., DC), chrominance / luminance component.

[0119] Figure 21 is a schematic diagram of a predictor / context selector unit 2100 that is similar in many respects to Figure 19. However, in Figure 21 the inputs (with respect to other coded / decoded code bits for which there are inputs) include inputs at relative locations A through F to the current code bit under consideration. The output may be a predicted value for the current code bit and / or a context for use in context-adaptive coding of the current code bit by the apparatus shown in Figure 15. For each data value in the set of data values ​​in the example samples A through F, one or more other data values ​​that predict the data value are located at predetermined relative locations to the data value in the array of data values.

[0120] In such an embodiment, encoding the data value codes may involve performing context-adaptive encoding (e.g., using Figure 15), where the context depends on the predicted data value codes (as obtained by the apparatus shown in Figure 21). Similarly, for decoding, the configuration shown in Figure 21 may be used to derive a context for use by the entropy decoder 410, which complements the apparatus shown in Figure 15. Thus, decoding the data value codes involves performing context-adaptive decoding, where the context depends on the predicted data value codes.

[0121] 21 and 19 may be, for example, the signs of other data values. In this case, the properties of the one or more data values ​​include the signs of the one or more other data values. In another example, the properties of the one or more other data values ​​may include the magnitudes of the one or more other data values.

[0122] As a summary of the information presented above, FIG. 22 is a schematic flow chart of a data encoding method, which includes: Encoding an ordered array of data values ​​as data representing the magnitude of the data values ​​and data representing the code of the data values ​​(step 2200); predicting a code for each data value from properties of one or more other data values ​​in the ordered array for a set of data values ​​that includes at least some of the data values ​​(step 2210); Encoding data value codes for the set of data values ​​based on each predicted value (step 2220).

[0123] Similarly, FIG. 23 is a schematic flow chart showing a data decoding method, which includes: Decoding the ordered array of data values ​​as data representing the magnitude of the data values ​​and data representing the code of the data values ​​(step 2300); predicting a code for each data value from properties of one or more other data values ​​in the ordered array for a set of data values ​​that includes at least some of the data values ​​(step 2310); Decoding data value codes for the set of data values ​​based on each predicted value (step 2320).

[0124] FIG. 24 is a schematic diagram illustrating at least a portion of a data encoding apparatus, the data encoding apparatus including: an encoder 2400 configured to encode an ordered array of data values ​​as data representing the magnitude of the data values ​​and data representing the code of the data values, the encoder 2400 including, for example, units 360, 370 as shown in FIG. 7; a predictor 2410, e.g., implemented by unit 370 and unit 343 working together, configured to predict a code of each data value from properties of one or more other data values ​​in the ordered array for a set of data values ​​including at least some of the data values; The encoder is configured to encode a data value code for the set of data values ​​based on each predictor value.

[0125] FIG. 24 is a schematic diagram illustrating at least a portion of a data decoding apparatus, the data decoding apparatus including: a decoder 2500 configured to encode an ordered array of data values ​​as data representing the magnitude of the data values ​​and data representing the code of the data values, the decoder 2500 including, for example, units 410, 400 shown in FIG. 7; a predictor 2510, e.g., implemented by unit 410 and unit 343 working together, configured to predict a code of each data value from properties of one or more other data values ​​in the ordered array for a set of data values ​​including at least some of the data values; The decoder includes encoding a data value code for the set of data values ​​based on each predictor value.

[0126] One or both of the encoding and decoding devices may be implemented as at least part of a video storage, capture, transmission, or reception device.

[0127] One example of a coding order has already been described above. Two other examples are described below with reference to Figures 26 and 27. In Figures 26 and 27, the left column shows the operations performed for a transform unit, and the right column shows the operations performed for a 4x4 or other group of 16 (or other subdivision) coefficients in the transform unit. Each of these is processed in turn, for example, according to a reverse diagonal scan order for the 16 coefficient groups.

[0128] Coding order example Referring to FIG. 26, the conversion unit can perform the following steps: Coding of the last X / Y position, which is the X / Y position of the coefficient with the highest scan order position (highest frequency coefficient) (step 2600) For each group of 16 coefficients, start with the last coefficient group that contains the last X / Y position. Coding whether the group contains a coefficient (if not already predicted by other methods) (step 2610) · Code the validity map (step 2620) Coding two or more (>1) maps (step 2630) Coding 3 or more (>2) maps (step 2640) Coding the remaining coefficients (step 2650) If the number of coefficients in the group exceeds the limit, signal the EMT+NSST flag (if necessary) (step 2660) For the 16 coefficients in each group, · Code the sign bit (step 2670)

[0129] Referring to FIG. 27, in another configuration example, the conversion unit can perform the following steps. Coding of the last X / Y position, which is the X / Y position of the coefficient with the highest scan order position (highest frequency coefficient) (step 2700) If the last X / Y position (which is the index in the scan order) is out of bounds, signal the EMT+NSST flag (if necessary) (step 2710) For each group of 16 coefficients, start with the last coefficient group that contains the last X / Y position. Coding whether the group contains a coefficient (if not already predicted by other methods) (step 2720) · Code the validity map (step 2730) Coding two or more (>1) maps (step 2740) Coding 3 or more (>2) maps (step 2750) Coding the remaining coefficients (step 2760) · Coding the sign bit (step 2770)

[0130] Steganography and Sign Bit Hiding (SBH) This means hiding data within other data, and in this context refers to a technique known as Sign Bit Hiding (SBH), a version of which is applicable to the present technology.

[0131] Sign bit hiding is a technique used to reduce the cost of transmitting one sign bit for each block by (effectively) hiding the sign bits of non-zero coefficients within other groups of coefficients. This is done by the encoder artificially setting the parity (whether the sum is even or odd) of the coefficient groups to a desired value, so that the parity itself indicates the hidden sign bit. This can be achieved by slightly distorting the coefficient values, preferably in such a way that the increased noise does not outweigh the net gain of not transmitting one sign bit.

[0132] The SBH technique described above applies this to the DC coefficient or the last coefficient to be coded according to coding or array order (described below).

[0133] In the example of the present technology, SBH or steganographic coding is applied to the first code bit to be coded in the array order, because the prediction of the code bit mentioned above depends on the code bit that has already been coded or decoded. Therefore, a larger gain can be obtained by using SBH for the first code bit to be decoded, so that other predictions can also be made based on this code bit.

[0134] The array processing order for encoding the codes may be, for example, a reverse diagonal scan order as shown schematically in Figure 28. The first code bit to encode is the sign bit of the coefficient in the "last X / Y position" in the current group of 16 coefficients, for example, coefficient 2800 in Figure 28. The "last X / Y position" is derived from the top left corner of the block, but encoding and decoding are performed according to the array order shown in Figure 28.

[0135] Thus, in these examples, the method includes predicting and encoding data values ​​in the array according to an array processing order, such as the reverse diagonal scan order shown in Figure 28. In some examples, the method includes steganographically encoding a data value code of the first data value according to the array processing order (e.g., the "last X / Y position" other than the reverse diagonal order shown in Figure 28) in data representing the magnitude of the data values.

[0136] To the extent that embodiments of the present invention are described as being implemented, at least in part, by a software-controlled data processing apparatus, it will be understood that a non-transitory machine-readable medium, such as an optical disk, magnetic disk, semiconductor memory, or the like, having such software thereon is also considered to represent an embodiment of the present disclosure. Moreover, such software may be distributed in other forms. Similarly, a data signal (whether embodied in a non-transitory machine-readable medium or not) containing encoded data produced according to the method described above is also considered to represent an embodiment of the present disclosure.

[0137] Numerous modifications and variations of the present disclosure are possible in light of the above teachings, and it is therefore to be understood that within the scope of the appended claims, the technology may be practiced otherwise than as specifically described herein.

[0138] It should be noted that in the above description, for clarity of explanation, different embodiments have been described with reference to different functional units, circuits and / or processors, but it will be apparent that any suitable division of roles between different functional units, circuits and / or processors can be adopted without adversely affecting the embodiments.

[0139] The above embodiments may be implemented in any suitable manner, including hardware, software, firmware, or any combination thereof. The above embodiments may optionally be implemented at least in part as computer software running on one or more data processors and / or digital signal processors. The elements and components of any embodiment may be physically, functionally, or logically implemented in any manner. Indeed, the functionality may be implemented in a single unit, in multiple units, or as part of other functional units. Thus, the disclosed embodiments may be implemented in a single unit or may be physically or functionally distributed between different units, circuits, and / or processors.

[0140] While the present disclosure has been described in connection with certain embodiments, it is not intended to be limited to those specifically recited, and while certain features may appear to be described in connection with particular embodiments, those skilled in the art will recognize that various features of the disclosed embodiments may be combined in any suitable manner and used to practice the techniques.

[0141] Each aspect and feature of the present disclosure is defined by the following items. 1. A data encoding method comprising: encoding an ordered array of data values ​​as data representing magnitudes of the data values ​​and data representing codes of the data values; predicting a code for each data value from properties of one or more other data values ​​in the ordered array for a set of data values ​​that includes at least some of the data values; encoding a data value code for the set of data values ​​based on each predicted value; Data encoding method. 2. The method according to item 1, the step of encoding the data value codes includes performing context adaptive encoding; The context depends on the expected data value sign, method. 3. The method according to item 1 or 2, the properties of the one or more other data values ​​include the signs of the one or more other data values; method. 4. A method according to any one of the above items, comprising: the properties of the one or more other data values ​​include the magnitude of said one or more other data values; method. 5. A method according to any one of the above items, comprising: for each data value in the set of data values, the one or more other data values ​​that predict the data value are located at a predetermined relative position to the data value in the array of data values; method. 6. The method according to item 5, For each data value in the set of data values, the relative position of the one or more other data values ​​that predicts the data value depends on the position of the predicting data value within the array of data values. method. 7. The method according to item 5 or 6, The data values ​​represent image data. method. 8. The method according to item 7, predicting an image region of an image to be coded; generating a residual image region dependent on the difference between a predicted image region and a corresponding region of the image to be coded, the ordered array of data values ​​comprises data values ​​of the representation of the residual image region; method. 9. The method according to item 8, the ordered array of data values ​​comprises data values ​​of a frequency transform representation of the residual image region; the residual image region is a region that has already undergone a series of one or more frequency transforms; method. 10. The method according to item 9, The frequency conversion is Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), DCT in one direction and DST in the orthogonal direction, and A linear transformation followed by a non-separable quadratic transformation, including one or more of method. 11. The method according to item 8, the ordered array of data values ​​comprises data values ​​of a reordered representation of the residual image region in transform skip mode; method. 12. The method according to any one of items 8 to 11, the step of predicting an image region of the image to be coded comprises predicting a sample of the image region based on other coded and decoded samples of the image displaced from the predicted sample in a direction defined by the prediction mode; method. 13. The method according to item 12, For each data value in the set of data values, the relative position of the one or more other data values ​​from which the data value is predicted depends on a prediction mode applicable to the array of data values. method. 14. The method according to item 12 or 13, For each data value in the set of data values, the properties of one or more other data values ​​that predict the data value are: a prediction mode applicable to the data value array; The array size of the data value array, The shape of the data value array, and The location of the data value within the data value array Depends on one or more elements selected from the list consisting of method. 15. A method according to any one of the preceding items, comprising: predicting and encoding data values ​​of the array in accordance with an array processing order; method. 16. The method according to item 15, encoding a data value code of a first data value among data representing the magnitude of the data value using steganography in accordance with an array processing order; method. 17. When executed by a computer, causes the computer to perform the method described in any one of the preceding items; Computer software. 18. A machine-readable non-transitory storage medium storing the computer software described in item 17. 19. A data encoding device, comprising: an encoder configured to encode an ordered array of data values ​​as data representing the magnitude of the data values ​​and data representing the code of the data values; a predictor configured to predict, for a set of data values ​​including at least some of the data values, a code for each data value from a property of one or more other data values ​​in the ordered array; the encoder is configured to encode data value codes for the set of data values ​​based on each predicted value. Data encoding device. 20. A device according to item 19, Video storage, capture, transmission or receiving equipment. 21. A data decoding method comprising: decoding the ordered array of data values ​​as data representing the magnitude of the data values ​​and data representing the code of the data values; predicting a code for each data value from properties of one or more other data values ​​in the ordered array for a set of data values ​​that includes at least some of the data values; decoding data value codes for the set of data values ​​based on each predicted value; Data decryption method. 22. The method according to item 21, the step of decoding the data value code includes performing context adaptive decoding; The context depends on the expected data value sign, method. 23. The method according to item 21 or 22, the properties of the one or more other data values ​​include the signs of the one or more other data values; method. 24. The method according to any one of items 21 to 23, the properties of the one or more other data values ​​include the magnitude of said one or more other data values; method. 25. The method according to any one of items 21 to 24, for each data value in the set of data values, the one or more other data values ​​that predict the data value are located at a predetermined relative position to the data value in the array of data values; method. 26. The method according to item 25, For each data value in the set of data values, the relative position of the one or more other data values ​​that predicts the data value depends on the position of the predicting data value within the array of data values. method. 27. The method according to item 25 or 26, The data values ​​represent image data. method. 28. The method according to item 27, predicting an image region of an image to be decoded; generating residual image regions indicative of differences between predicted image regions and corresponding regions of the image to be decoded; the ordered array of data values ​​comprises data values ​​of a representation of the residual image region; The method further comprises combining the predicted image region and the residual image region; method. 29. The method according to item 28, the ordered array of data values ​​comprises data values ​​of a frequency transform representation of the residual image region; method. 30. The method according to Item 29, The frequency conversion is Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), DCT in one direction and DST in the orthogonal direction, and A linear transformation followed by a non-separable quadratic transformation, Including, method. 31. The method according to item 28, the ordered array of data values ​​comprises data values ​​of a reordered representation of the residual image region in transform skip mode; method. 32. The method according to any one of items 28 to 31, the step of predicting an image region of the image to be decoded comprises predicting a sample of the image region based on other decoded samples of the image displaced from the predicted sample in a direction defined by the prediction mode; method. 33. The method according to item 32, For each data value in the set of data values, the relative position of the one or more other data values ​​from which the data value is predicted depends on a prediction mode applicable to the array of data values. method. 34. The method according to item 32 or 33, For each data value in the set of data values, the properties of one or more other data values ​​that predict the data value are: a prediction mode applicable to the data value array; The array size of the data value array, The shape of the data value array, and The location of the data value within the data value array Depends on one or more elements selected from the list consisting of method. 35. The method according to any one of items 21 to 34, predicting and decoding data values ​​of the array in accordance with an array processing order; method. 36. The method according to item 35, encoding a data value code of a first data value among data representing the magnitude of the data value using steganography in accordance with an array processing order; method. 37. When executed by a computer, causes the computer to perform the method described in item 21. Computer software. 38. A machine-readable non-transitory storage medium storing the computer software described in item 37. 39. A data decoding device, comprising: a decoder configured to encode an ordered array of data values ​​as data representing the magnitude of the data values ​​and data representing the code of the data values; a predictor configured to predict, for a set of data values ​​including at least some of the data values, a code of each data value from a property of one or more other data values ​​in the ordered array; the decoder is configured to encode data value codes for the set of data values ​​based on each predicted value; Data decoding device. 40. A device comprising the device according to item 39. Video storage, capture, transmission or receiving equipment. 41. A video capture device includes an image sensor and an encoding device according to item 19. 42. The video capture device according to item 41 further includes a display to which the device according to item 39 and the data stream are output. 43. The video imaging device according to item 41 includes a transmitting unit configured to transmit the encoded data stream.

Claims

1. 1. A data encoding method comprising: encoding an ordered array of data values ​​in two or more modes, wherein in a first mode the ordered array of data values ​​represents DCT coefficients and in a second mode, which is a transform skip mode, the ordered array of data values ​​represents data value magnitudes and data value signs; predicting, for a set of data values ​​including at least a plurality of data values ​​in the transform skip mode, a code for each data value from a plurality of adjacent data values ​​in the ordered array from previously encoded data value codes; encoding a data value signature dependent on each predicted data value signature for the set of data values; obtaining a context for encoding a data value in the transform skip mode, the context including a data value code; encoding a data value code using the obtained context; Including, the ordered array of data values ​​comprises data values ​​of a reordered representation of a residual image region in the transform skip mode. Data encoding method.

2. 1. A data encoding method comprising: encoding an ordered array of data values ​​in two or more modes, the data values ​​representing DCT or DST coefficients in a first mode and data values ​​representing data value magnitudes and data value signs in a second mode, which is a transform skip mode; predicting, for data values ​​in the second mode, each data signature value from one or more data values ​​in the ordered array; obtaining a context for encoding a data value; encoding a data value code using the obtained context; Including, the ordered array of data values ​​comprises data values ​​of a reordered representation of a residual image region in the transform skip mode. Data encoding method.

3. 3. The data encoding method according to claim 2, The context depends on each detected data code value. Data encoding method.

4. 3. The data encoding method according to claim 2, detecting each data code value from one or more data values ​​in the ordered array includes detecting data code values ​​for neighboring already-encoded data values. Data encoding method.

5. 3. The data encoding method according to claim 2, detecting each data code value from one or more data values ​​in the ordered array includes detecting data code values ​​for horizontally adjacent data values ​​and vertically adjacent data values. Data encoding method.

6. 5. The data encoding method according to claim 4, obtaining the context includes using a lookup function of the detected data signature value; Data encoding method.

7. 7. A data encoding method according to claim 6, comprising: the lookup function determines which of a plurality of contexts to select based on whether the detected data code values ​​are different or whether the detected data code values ​​are the same; encoding the data value based on the selected context; Data encoding method.

8. 8. A data encoding method according to claim 7, comprising: the lookup function for determining which of a plurality of contexts to select further depends on the detected prediction mode. Data encoding method.

9. 9. A data encoding method according to claim 8, comprising: the detected prediction mode includes an inter prediction mode and an intra prediction mode; Data encoding method.

10. 2. The data encoding method according to claim 1, The context is a CABAC context. Data encoding method.

11. 1. A data encoding device, comprising: encoding an ordered array of data values ​​in two or more modes, the data values ​​representing DCT or DST coefficients in a first mode and data values ​​representing data value magnitudes and data value signs in a second mode, which is a transform skip mode; predicting, for data values ​​in the second mode, each data signature value from one or more data values ​​in the ordered array; obtaining a context for encoding a data value; encoding a data value code using the obtained context; a circuit configured to perform the the ordered array of data values ​​comprises data values ​​of a reordered representation of a residual image region in the transform skip mode. Data encoding device.

12. 1. A method for decoding data, comprising: decoding an ordered array of data values ​​in two or more modes, wherein in a first mode the ordered array of data values ​​represents DCT or DST coefficients and in a second mode, which is a transform skip mode, the ordered array of data values ​​represents data value magnitudes and data value signs; predicting, for data values ​​in the second mode, each data signature value from one or more data values ​​in the ordered array; obtaining a context for encoding a data value; Decoding the data value code using the obtained context. Including, the ordered array of data values ​​comprises data values ​​of a reordered representation of a residual image region in the transform skip mode. Data decoding method.

13. 13. A data decoding method according to claim 12, comprising: The context depends on each detected data code value. Data decoding method.

14. 13. A data decoding method according to claim 12, comprising: detecting each data code value from one or more data values ​​in the ordered array includes detecting data code values ​​for neighboring already-encoded data values. Data decoding method.

15. 13. A data decoding method according to claim 12, comprising: detecting each data code value from one or more data values ​​in the ordered array includes detecting data code values ​​for horizontally adjacent data values ​​and vertically adjacent data values. Data decoding method.

16. 16. A data decoding method according to claim 15, comprising: obtaining the context includes using a lookup function of the detected data signature value; Data decoding method.

17. 17. A method for decoding data according to claim 16, comprising: the lookup function determines which of a plurality of contexts to select based on whether the detected data code values ​​are different or whether the detected data code values ​​are the same; and decoding the data value based on the selected context. Data decoding method.

18. 18. A method for decoding data according to claim 17, comprising: the lookup function for determining which of a plurality of contexts to select further depends on the detected prediction mode. Data decoding method.

19. 20. A method for decoding data according to claim 18, comprising: the detected prediction mode includes an inter prediction mode and an intra prediction mode; Data decoding method.

20. 13. A data decoding method according to claim 12, comprising: The context is a CABAC context. Data decoding method.

21. A computer-readable recording medium having recorded thereon a software program that, when executed by a computer, causes the computer to perform the data decoding method of claim 12.

22. 1. A data decoding device, comprising: decoding an ordered array of data values ​​in two or more modes, wherein in a first mode the ordered array of data values ​​represents DCT or DST coefficients and in a second mode, which is a transform skip mode, the ordered array of data values ​​represents data value magnitudes and data value signs; predicting, for data values ​​in the second mode, each data signature value from one or more data values ​​in the ordered array; obtaining a context for encoding a data value; Decoding the data value code using the obtained context. a circuit configured to perform the the ordered array of data values ​​comprises data values ​​of a reordered representation of a residual image region in the transform skip mode. Data decoding device.

23. An image processing device that functions as the data decoding device according to claim 22, Save, capture, transfer, or receive video images; Image processing device.

Citation Information

Patent Citations

  • Encoder and control method therefor

    JP2016092589A

  • Data Encoding and Decoding

    JP2016519514A

  • Coding sign information of video data

    WO2017083710A1