Image data encoding and decoding
The entropy encoder with attribute detection and padding ensures efficient data compression by managing byte size constraints, enhancing the efficiency of video data encoding and decoding processes.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SONY GROUP CORP
- Filing Date
- 2025-04-03
- Publication Date
- 2026-05-26
AI Technical Summary
Existing video data encoding systems face challenges in efficiently managing the byte size constraints of encoded data sections, leading to inefficiencies in data compression and decompression processes.
An entropy encoder that selectively encodes image data, detects encoding attributes, and applies constraints on the number of binarized symbols per byte, with padding data added to ensure byte size compliance, along with a decoding apparatus for decoding data signals generated by this method.
This approach enhances data compression efficiency by maintaining byte size constraints, thereby improving the overall data compression and decompression processes.
Smart Images

Figure 0007865414000002 
Figure 0007865414000003 
Figure 0007865414000004
Abstract
Description
Technical Field
[0001] The present disclosure relates to encoding and decoding of image data.
Background Art
[0002] The description of the "Background Art" in this specification is for generally explaining the background in the present application. The techniques of the present inventors should not be regarded as prior art at the time of filing of the present application if they are not prior art within the scope described in this background art section, whether explicitly or implicitly, as well as in terms of the aspect of the description that should not be regarded as prior art for the present application.
[0003] There are several video data encoding systems and video data decoding systems that convert video data into a frequency domain representation, quantize the obtained frequency domain coefficients, and then apply a certain type of entropy encoding to the quantized coefficients. Thereby, video data can be compressed. By applying the corresponding decoding or decompression technique, the original video data is reconstructed and restored.
[0004] As a successor technology to H.264 / MPEG-4 AVC, HEVC (High Efficiency Video Coding), also known as H.265 or MPEG-H Part 2, has been proposed. This aims to improve the video quality of HEVC and double the data compression ratio of H.264, and to make H.264 scalable within the range of 128×96 to 7680×4320 pixel resolutions. This corresponds to a bit rate of approximately 128 kilobits per second to 800 megabits per second.
Summary of the Invention
Problems to be Solved by the Invention
[0005] The present disclosure aims to address or mitigate the problems associated with the above processes.
Means for Solving the Problems
[0006] This disclosure provides an entropy encoder configured to selectively encode data items representing image data so as to generate encoded binarized symbols of consecutive output data portions, An attribute detector configured to detect an encoding attribute applicable to a given output data section, A selector configured to select a constraint from two or more constraint candidates for use with the given output data section, depending on the detected encoding attribute, It is equipped with, The above entropy encoder is configured to generate an output data stream, and this output data stream is subject to a constraint that defines an upper limit on the number of binarized symbols represented by the byte size of any individual output data section relative to the byte size of the output data section. To satisfy this constraint, the entropy encoder is configured to provide padding data to each output data section that does not satisfy the constraint, thereby increasing the byte-by-byte size of the output data section. We provide an image data encoding device.
[0007] This disclosure includes the steps of selectively encoding data items representing image data to generate encoded binarized symbols for consecutive output data portions, The steps include generating an output data stream that is subject to constraints defining an upper limit on the number of binarized symbols represented by any individual output data section with respect to the byte-unit size of the output data section, To satisfy this constraint, the step of providing padding data to each output data section that does not satisfy this constraint, in order to increase the byte size of the output data section, A step of detecting an encoding attribute applicable to a given output data section, The steps include: selecting a constraint from two or more candidate constraints to use with a given output data section in response to the detected coding attribute; including Further methods for encoding image data are provided.
[0008] This disclosure further provides suitable decoding apparatus for decoding data signals generated by the method or apparatus described above.
[0009] Further aspects and features of this disclosure are defined in the appended claims.
[0010] It should be understood that the general explanation above and the detailed explanation below are examples of this technology and do not limit it. [Brief explanation of the drawing]
[0011] By referring to the detailed explanation below, along with the attached drawings, a complete understanding of this technology and many of its advantages can be easily grasped. [Figure 1] This is a schematic diagram showing an audio / video (A / V) data transmission and reception system that compresses and decompresses video data. [Figure 2] This is a schematic diagram showing a video display system that performs video data decompression. [Figure 3] This is a schematic diagram showing an audio / video storage system that compresses and decompresses video data. [Figure 4] This is a schematic diagram showing a video camera that compresses video data. [Figure 5] This is a schematic diagram showing a storage medium. [Figure 6] This is a schematic diagram showing a storage medium. [Figure 7] This is a schematic diagram showing a video data compression and decompression device. [Figure 8] This is a schematic diagram showing the prediction unit. [Figure 9] This is a schematic diagram showing a partially encoded image. [Figure 10] This is a schematic diagram showing possible intra-prediction directions for the set. [Figure 11]It is a schematic diagram showing the prediction mode of the set. [Figure 12] It is a schematic diagram showing the prediction mode of another set. [Figure 13] It is a schematic diagram showing the intra prediction process. [Figure 14] It is a schematic diagram showing the CABAC encoder. [Figure 15] It is a schematic diagram showing the CABAC encoding method. [Figure 16] It is a schematic diagram showing the CABAC encoding method. [Figure 17] It is a schematic diagram showing the CABAC decoding technology. [Figure 18] It is a schematic diagram showing the CABAC decoding technology. [Figure 19] It is a schematic diagram showing the divided image. [Figure 20] It is a schematic diagram showing a device. [Figure 21] It is a schematic diagram showing the controller. [Figure 22] A and B schematically represent the output data part. [Figure 23] It is a schematic flowchart showing a method.
Embodiments for Carrying Out the Invention
[0012] Next, referring to each drawing, FIGS. 1 to 4 schematically show a device or a system that uses a compression device and / or a decompression device according to each embodiment of the present technology described below.
[0013] All data compression and / or decompression devices described below may be implemented in hardware, or they may be implemented as programmable hardware such as an Application Specific Integrated Circuit (ASIC) or a Field Programmable Gate Array (FPGA), or a combination thereof, in software running on a general-purpose data processing device such as a general-purpose computer. In embodiments implemented by software and / or firmware, it will be understood that such software and / or firmware, as well as non-temporary data recording media on which such software and / or firmware is stored or provided, are considered embodiments of the Technology.
[0014] Figure 1 is a schematic diagram showing an audio / video data transmission and reception system that compresses and decompresses video data.
[0015] The input audio / video signal 10 is supplied to a video data compression device 20 that compresses at least the video elements of the audio / video signal 10 and is transmitted along a transmission route 30, such as a cable, optical fiber, or wireless link. The compressed signal is processed by a decompression device 40, which provides the output audio / video signal 50. In the return path, a compression device 60 compresses the audio / video signal, and the audio / video signal is transmitted along the transmission route 30 to a decompression device 70.
[0016] Therefore, the compression device 20 and the decompression device 70 can constitute one node of the transmission link. The decompression device 40 and the compression device 60 can constitute another node of the same transmission link. Of course, if the transmission link is unidirectional, only one of these nodes will require a compression device, and only the other node will require a decompression device.
[0017] Figure 2 is a schematic diagram showing a video display system that performs video data decompression. Specifically, the compressed audio / video signal 100 is processed by the decompression device 110, thereby providing a decompressed signal that can be displayed on the display device 120. The defrosting device 110 may be integrally formed with the display device 120 by, for example, being installed in the same housing as the display device 120. Alternatively, the defrosting device 110 may be provided as a so-called set-top box (STB). It should be noted that the term "set-top" does not mean that the box must be positioned in a specific orientation or location relative to the display device 120. This term is simply used in the art to refer to a device that can be connected to the display unit as a peripheral device.
[0018] Figure 3 is a schematic diagram showing an audio / video storage system that compresses and decompresses video data. The input audio / video signal 130 is supplied to a compression device 140 that generates a compressed signal, and is stored in a storage device 150, such as a magnetic disk drive, optical disk drive, magnetic tape drive, or solid-state storage device such as semiconductor memory or other storage devices. During playback, the compressed data is read from the storage device 150 and sent to a decompression device 160 for decompression. This provides the output audio / video signal 170.
[0019] It will be understood that compressed or encoded signals and non-transient, device-readable storage media, etc., that store such signals are considered embodiments of this technology.
[0020] Figure 4 is a schematic diagram showing a video camera that compresses video data. In Figure 4, an image capture device 180, including a CCD (Charge Coupled Device) image sensor and associated control and readout electronic equipment, generates a video signal to be sent to a compression device 190. One or more microphones 200 generate an audio signal to be sent to the compression device 190. The compression device 190 generates a compressed audio / video signal 210 to be stored and / or transmitted (collectively represented as stage 220).
[0021] The techniques described below primarily relate to the compression and decompression of video data. It will be understood that, in order to compress audio data, many existing techniques may be used in conjunction with the video data compression techniques described below to generate compressed audio / video signals. Therefore, we will not provide a separate explanation of audio data compression. Furthermore, it will be understood that, particularly in broadcast-quality video data, the data rate associated with video data (whether compressed or uncompressed) is generally much higher than the data rate associated with audio data. Therefore, it will be understood that uncompressed audio data can be added to compressed video data to create a compressed audio / video signal. Furthermore, although embodiments of the present invention (see Figures 1 to 4) relate to audio / video data, it will be understood that the techniques described below may be used simply in systems that process (i.e., compress, decompress, store, display, and / or transmit) video data. In other words, these embodiments do not necessarily have to be related to audio data processing and can be applied to video data compression.
[0022] Therefore, Figure 4 provides an example of a video capture device including an image sensor and an encoding device of the type described below. Accordingly, Figure 2 provides an example of a decoding device of the type described below and a display device that outputs the decoded image.
[0023] The combination shown in Figures 2 and 4 can provide a video capture device that includes an image sensor 180, an encoding device 190, a decoding device 110, and a display device 120 on which the decoded image is output.
[0024] Figures 5 and 6 are schematic diagrams showing compressed data (for example) generated by devices 20 and 60, compressed data input to device 110, or storage media that store storage media, i.e., storage devices 150 and 220. Figure 5 is a schematic diagram showing a disk-type storage medium such as a magnetic disk or optical disk. Figure 6 is a schematic diagram showing a solid-state storage medium such as flash memory. Figures 5 and 6 also show examples of non-transient, device-readable storage media that store computer software that, when executed by a computer, causes the computer to execute one or more of the methods described later.
[0025] Therefore, the above configuration provides examples of video storage devices, capture devices, and transceivers that embody any of the present technologies.
[0026] Figure 7 is a schematic diagram of a video data compression and decompression device.
[0027] The control unit 343 controls the overall operation of the video data compression and decompression device. In particular, with respect to the compression mode, the control unit 343 controls the trial encoding process by acting as a selector that selects various operating modes such as block size and shape. Furthermore, the control unit 343 controls whether or not loss occurs when the video data is encoded. This control unit is considered to constitute part of the image encoder or image decoder (depending on the case). The continuous images of the input video signal 300 are supplied to the adder 310 and the image prediction unit 320. The image prediction unit 320 will be described in detail later with reference to Figure 8. The image encoder or image decoder (if applicable) may use features from the apparatus shown in Figure 7, together with the intra-image prediction unit in Figure 8. However, this image encoder or image decoder does not necessarily require all of the features shown in Figure 7.
[0028] The adder 310 receives the input video signal 300 on the "+" input and the output of the image prediction unit 320 on the "-" input, effectively performing a subtraction (negative addition) operation. As a result, the predicted image is subtracted from the input image. This generates a so-called residual image signal 330 that represents the difference between the actual image and the predicted image.
[0029] One reason for generating residual image signals is as follows: The data encoding techniques used to explain residual image signals tend to work more efficiently when the image being encoded has less "energy." Here, "efficient" means that the amount of encoded data generated is small. At certain image quality levels, it is desirable (and considered "efficient") to generate as little data as possible. The "energy" in a residual image is related to the amount of information contained in the residual image. If the predicted image and the actual image are identical, the difference between these two images (i.e., the residual image) contains zero information (zero energy) and can be very easily encoded into a small amount of coded data. Generally, if the prediction process can be performed reasonably well so that the content of the predicted image is similar to the content of the encoded image, the residual image data is expected to contain less information (less energy) than the input image and can be easily encoded into a small amount of encoded data.
[0030] The remaining part of the device, which acts as an encoder (encoding residual or difference images), is described below. The residual image data 330 is supplied to a transform unit, i.e., circuit 340, which generates a discrete cosine transform (DCT) representation of blocks or regions of the residual image data. The DCT technique itself is widely known and will not be described in detail here. Furthermore, the use of DCT is merely an example of one possible configuration. Other transformation methods could be used, for example, the Discrete Sine Transform (DST). The transformation methods may also be combined, for example, in a configuration where one transformation method is followed by another (directly or indirectly). The selection of the conversion method may be explicitly determined, or / or depend on the side information used to configure the encoder and decoder.
[0031] The output of the transformation unit 340, i.e., a series of DCT coefficients for each transformation block in the image data, is supplied to the quantization unit 350. Various quantization techniques are widely known in the field of video data compression, ranging from simple multiplication by quantization scaling elements to the application of complex lookup tables under the control of quantization parameters. There are two common purposes for this. The first is to reduce the range of values that the transformation data can take through the quantization process. The second is to increase the probability that the value of the converted data is zero through quantization. These factors allow for more efficient entropy coding, which will be discussed later, when generating small amounts of compressed video data.
[0032] The scanning unit 360 applies data scanning processing. The purpose of the scanning processing is to reorganize the quantized data in order to group non-zero quantization conversion coefficients together as much as possible, and of course, to group zero-value coefficients together as much as possible. These functions allow for the efficient application of so-called run-length coding or similar techniques. Therefore, the scanning process includes selecting coefficients from blocks of coefficients corresponding to blocks of quantized and quantized image data, according to a “scan order,” such that (a) all coefficients are selected at least once as part of the scan, and (b) the scan can perform the desired reorganization. An example of a scan order that yields valid results is the so-called Up-right Diagonal scan order.
[0033] The scanned coefficients are then sent to the entropy encoder (EE) 370. In this case, too, various entropy coding methods may be performed. Two examples are a variation of the so-called CABAC (Context Adaptive Binary Coding) system and a variation of the so-called CAVLC (Context Adaptive Variable-Length Coding) system. Generally, CABAC is considered efficient. One study showed that the amount of coded output data in CABAC is 10-20% less than in CAVLC for equivalent image quality. However, the level of complexity (in execution) exhibited by CAVLC is considered to be far lower than that of CABAC. Although the scanning and entropy coding processes are presented as separate operations, in practice they can be combined or handled together. That is, data can be read into the entropy encoder in scan order. This is also true for the inverse operations described later.
[0034] The output of the entropy encoder 370 provides a compressed output video signal 380, along with additional data (as described above and / or below) that defines, for example, how the prediction unit 320 generates the predicted image.
[0035] On the other hand, since the operation of the prediction unit 320 itself depends on the decompressed compressed output data, a return path is also provided.
[0036] The reason for this function is as follows: Decompressed residual data is generated at an appropriate stage in the decompression process (described later). This decompressed residual data needs to be added to the predicted image in order to generate the output image (because the original residual data was the difference between the input image and the predicted image). In order for this process to be equivalent on both the compression and decompression sides, the predicted image generated by the prediction unit 320 should be identical during both the compression and decompression processes. Of course, the device cannot access the original input image during decompression. The device can only access the decompressed image. Therefore, during compression, the prediction unit 320 makes predictions (at least for inter-image coding) based on the decompressed compressed image.
[0037] The entropy encoding process performed by the entropy encoder 370 can be considered "lossless" (at least in some examples). That is, it can be replaced with data that is exactly the same as the data initially supplied to the entropy encoder 370. Therefore, in such examples, the return path can be implemented before the entropy encoding stage. In fact, the scan process performed by the scan unit 360 is also considered to be lossless, but in this embodiment, the return path 390 extends from the output of the quantization unit 350 to the input of the supplemental inverse quantization unit 420. If a stage causes or may cause loss, that stage may be included in the feedback loop formed by the return path. For example, the entropy coding stage may be designed to produce loss, at least in principle, by techniques such as coding bits in parity information. In such cases, entropy coding and decoding need to form part of a feedback loop.
[0038] Generally, the entropy decoder 410, the inverse scan unit 400, the inverse quantization unit 420, and the inverse transform unit, i.e., the circuit 430, provide the inverse functions of the entropy encoder 370, the scan unit 360, the quantization unit 350, and the transform unit 340, respectively. Here, we will continue to explain the compression process, and the process for decompressing the input compressed video signal will be described separately later.
[0039] In the compression process, the scanned coefficients are sent from the quantization unit 350 via the return path 390 to the inverse quantization unit 420, which performs the reverse operation of the scanning unit 360. The inverse quantization and inverse transformation processes are performed by the inverse quantization unit 420 and the inverse transformation unit 430, and the compressed-decompressed residual image signal 440 is generated.
[0040] The image signal 440 is added to the output of the prediction unit 320 by the summing unit 450, generating the reconstructed output image 460. This constitutes one input to the image prediction unit 320, as will be described later.
[0041] The process applied to decompress the received compressed video signal 470 is described below. The compressed video signal 470 is first supplied to the entropy decoder 410, from which it is supplied in the order of the inverse scan unit 400, the inverse quantization unit 420, and the inverse transform unit 430. It is then added to the output of the image prediction unit 320 by the adder unit 450. Therefore, on the decoder side, the decoder reconstructs the residual image and decodes each block by applying it (block by block) to the predicted image (by the summer 450). In short, the output 460 of the summer 450 forms the output decompressed video signal 480. In practice, further filtering may be optionally applied before outputting the signal (for example, using filter 560). This filter 560 is shown in Figure 8. In Figure 7, which shows the overall configuration compared to Figure 8, filter 560 is omitted for clarity.
[0042] The devices shown in Figures 7 and 8 can operate as a compression (encoding) device or a decompression (decoding) device. The functions of the two types of devices substantially overlap. The scan unit 360 and the entropy encoder 370 are not used in decompression mode. The prediction unit 320 (described in detail later) and the other units operate according to the mode and parameter information contained in the received compressed bitstream, and do not generate this information themselves.
[0043] Figure 8 is a schematic diagram illustrating the generation of a predicted image, and in particular, shows the operation of the image prediction unit 320.
[0044] Two basic prediction modes are performed by the image prediction unit 320. These two basic prediction modes are so-called intra-image prediction and so-called inter-image prediction or motion-compensated (MC) prediction. On the encoder side, each of these predictions includes detecting the prediction direction for the current block to be predicted and generating a predicted block of the sample in accordance with other samples (in the same (intra) or different (inter) image). The adder unit 310 or 450 encodes or decodes the blocks by encoding or decoding the difference between the predicted block and the actual block.
[0045] (On the decoder side or the reverse decoding side of the encoder, this detection of the predicted direction may correspond to data associated with the encoded data by the encoder, indicating which direction was used by the encoder. Alternatively, the detection of the predicted direction may correspond to the same element determined by the encoder.)
[0046] Intra-image prediction is based on predicting the content of image blocks or regions in data obtained from within the same image. This corresponds to so-called I-frame coding in other video compression techniques. However, in contrast to I-frame coding, which encodes the entire image using intra-coding, this embodiment allows the selection between intra-coding and inter-coding on a block-by-block basis. In other embodiments, this selection is still made on an image-by-image basis.
[0047] Motion compensation prediction is an example of inter-image prediction, where motion information in other adjacent or nearby images is used to define the source of image detail encoded in the current image. Therefore, in an ideal example, the contents of blocks of image data in the predicted image can be very easily encoded as a reference (motion vector) that points to a corresponding block located at the same or slightly different position in the adjacent image.
[0048] The technique known as "block copy" prediction is, in some ways, a hybrid of the two prediction methods described above, as it uses a vector indicating a block consisting of samples located at a displaced position from the current predicted block within the same image, which should be copied to generate the current predicted block.
[0049] Returning to Figure 8, Figure 8 shows two image prediction configurations (corresponding to intra-image prediction and inter-image prediction), the prediction results of which are selected by the multiplier 500 under the control of the mode signal 510 (for example, of the control unit 343) to provide blocks of predicted images to be supplied to the adders 310 and 450. This selection is based on which option will result in the least "energy" (which can be thought of as the amount of information that needs to be encoded, as described above), and this selection is communicated to the decoder in the encoded output data stream. In this regard, for example, the image energy can be detected by trial subtracting regions of two versions of the predicted image from the input image, squaring each pixel value in the difference image, summing the multipliers, and identifying which of the two versions has a lower average multiplier for the difference image associated with that image region. In another example, trial coding can be performed for each selection or for each possible selection. The selection is then made according to the cost for each possible selection, relating to either or both of the number of bits required for coding and the distortion to the image.
[0050] In the intra-prediction system, the actual prediction is based on image blocks received as part of signal 460. That is, the prediction is based on encoded-decoded image blocks so that the exact same prediction can be made in the decompression device. However, the operation of the intra-image prediction unit 530 can also be controlled by the intra-mode selection unit 520 by deriving data from the input video signal 300.
[0051] In inter-image prediction, the motion compensation (MC) prediction unit 540 uses motion information such as motion vectors derived from the input video signal 300 by the motion estimation unit 550. The motion compensation prediction unit 540 applies these motion vectors to the reconstructed image 460 to generate inter-image prediction blocks.
[0052] Therefore, the intra-image prediction unit 530 and the motion compensation prediction unit 540 (which operates together with the estimation unit 550) operate as detection units that detect the prediction direction for the current block that is the target of prediction, and as generation units that generate predicted blocks of samples (which form part of the prediction results sent to the addition units 310 and 450) according to other samples defined by the prediction direction.
[0053] Here, we will describe the processing applied to signal 460. First, signal 460 is optionally filtered by the filter unit 560. The filter unit 560 will be described in more detail below. In this process, a "deblocking" filter is applied to eliminate, or at least mitigate, the impact on block-based processing and subsequent operations performed by the conversion unit 340. It is also possible to use a Sample Adaptive Offsetting (SAO) filter. Alternatively, an adaptive loop filter can be optionally applied using coefficients obtained by processing the reconstructed signal 460 and the input video signal 300. This adaptive loop filter is a type of filter that applies adaptive filter coefficients to the data to be filtered using known techniques. That is, the filter coefficients can vary based on various factors. Data defining which filter coefficients to use is inserted into a portion of the encoded output data stream.
[0054] When the device is operating as a decompression device, the filtered output from the filter unit 560 actually forms the output video signal 480. This signal is stored in one or more image or frame storage units 570. The storage of continuous images is necessary for motion compensation prediction processing, in particular for generating motion vectors. To secure the necessary memory, the images stored in the image storage unit 570 may be held in a compressed format and then decompressed for use in generating motion vectors. For this particular purpose, any known compression / decompression system may be used. The stored image is sent to an interpolation filter 580 that generates a higher-resolution stored image. In this example, the interpolation filter 580 generates intermediate samples (subsamples) such that the resolution of the interpolated image output is four times (in each dimension) that of the image stored in the image storage unit 570 when the luminance channel is 4:2:0, and eight times (in each dimension) that of the image stored in the image storage unit 570 when the color channel is 4:2:0. The interpolated image is sent as input to the motion estimation unit 550 and the motion compensation prediction unit 540.
[0055] Here, we will describe a method for dividing an image for compression. At a basic level, the image to be compressed can be thought of as an array of blocks or regions consisting of samples. Dividing an image into such blocks or regions can be done using a decision tree, as described in Bross et al., "High Efficiency Video Coding (HEVC) text specification draft 6," JCTVC-H1003_d0 (November 2011), the contents of which are incorporated herein by reference. In some examples, the resulting blocks or regions can be of various sizes and, in some cases, have a shape that, by decision tree, follows the overall arrangement of image features within the image. This alone can improve coding efficiency because samples representing or following similar image features tend to be grouped by such a configuration. In some examples, square blocks or regions of different sizes (e.g., 4x4 samples to, for example, 64x64, or larger blocks) are available for selection. Other configurations may use blocks or regions of different shapes, such as rectangular blocks (for example, oriented vertically or horizontally). Other non-square and non-rectangular blocks are also included. As a result of dividing the image into such blocks or regions, (at least in this example) each sample of the image is assigned to one, or even just one, block or region sample.
[0056] Next, we will explain the intra-prediction process. Generally, intra-prediction involves generating a prediction result for the current block of samples from previously encoded and decoded samples within the same image.
[0057] Figure 9 is a schematic diagram showing a partially encoded image 800. Here, the image is encoded in blocks from the upper left to the lower right. Block 810 is an example of a block that is in the process of being encoded during the processing of the entire image. The area from the shaded region 820 at the top to the left of block 810 has already been encoded. Any part of the shaded region 820 can be used for intra-image prediction of the contents of block 810, but the unshaded region below it cannot be used.
[0058] In some examples, the image is encoded in blocks such that larger blocks (referred to as Coding Units (CUs)) are encoded in an order such as the one described with reference to Figure 9. Each CU may be processed as two or more smaller blocks or Transform Units (TUs) in a set (depending on the block division process performed). This gives a hierarchical encoding order such that the image is encoded in units of CUs. Each CU is potentially encoded at the TU level. The above-mentioned hierarchical encoding order (per CU, then per TU) for each TU (the largest node in the block-partitioned tree structure) within the current CTU (Coding Tree Unit) means that there are previously encoded samples in the current CU that may be available for encoding that TU. These samples might be located, for example, in the upper right or lower left of the TU.
[0059] Block 810 represents a CU. As mentioned above, for intra-image prediction processing, this may be subdivided into smaller units of the set. An example of the current TU830 is shown within CU810. More generally, an image is divided into regions or sample groups so that signaling information and transformed data can be efficiently encoded. The signaling of information may require different tree structures consisting of subdivisions of the structure of the transformation, and in fact of the structure of the prediction information or the prediction itself. For these reasons, the CU may have different tree structures for the transformation blocks or regions, prediction blocks or regions, and prediction information. In some examples, such as HEVC, this structure can be a so-called quadtree of the CU, where leaf nodes contain one or more prediction units and one or more TUs. The TU may contain multiple transformation blocks corresponding to the luma and chroma representations of the image, and the prediction scheme can be considered applicable at the transformation block level. In some examples, the parameters applied to a particular sample group can be considered to be defined primarily at the block level. This block level may not have the same granularity as the transformation structure.
[0060] Intra-image prediction considers encoded samples before considering the current TU. These samples are those above and / or to the left of the current TU. The samples from which the required samples are predicted may be located in different positions or directions relative to the current TU. To determine which direction is suitable for the current prediction unit, the mode selection unit 520 of the exemplary encoder can try all available combinations of TU structures for each candidate direction and select the prediction direction and TU structure that yields the highest compression efficiency.
[0061] The image may be encoded slice by slice. In one example, a slice is a group of horizontally adjacent CUs. However, more generally, a slice can be made up of the entire residual image, or a slice can be a single CU or a row of CUs, etc. Since slices are encoded as independent units, some degree of error resilience is obtained. The encoder and decoder states are completely reset at slice boundaries. For example, intra-prediction is not performed across slice boundaries. For this reason, slice boundaries are treated as image boundaries.
[0062] Figure 10 is a schematic diagram showing the predicted directions of possible (candidate) sets. All direction candidates are available for the prediction unit. The direction is determined by horizontal and vertical movement relative to the current block position, but is encoded as a prediction "mode". The directions for the set are shown in Figure 11. Note that the so-called DC mode represents the simple arithmetic mean of the surrounding upper and left samples. Also, the set of directions shown in Figure 10 is just one example. In other examples, as schematically shown in Figure 12, one set may consist of (for example) 65 angular modes combined with DC and planar (a total of 67 modes). Other numbers of modes are also possible.
[0063] Generally, the system can operate to generate predicted blocks of samples based on other samples determined by the predicted direction after detecting the predicted direction. In some examples, the image encoder is configured to encode data that identifies the selected predicted direction for each sample or region of the image (and the image decoder is configured to detect such data).
[0064] Figure 13 is a schematic diagram illustrating the intra-prediction process. In this intra-prediction process, a sample 900 of a block or region 910 consisting of samples is derived from other reference samples 920 of the same image according to a direction 930 determined by the intra-prediction mode associated with the sample. In this example, the reference samples 920 are based on the blocks above and to the left of the target block 910, and the predicted value of sample 900 is obtained by tracking the reference samples 920 along the direction 930. Direction 930 may indicate a single, individual reference sample, but more generally, the interpolated value of surrounding reference samples is used as the predicted value. Note that block 910 may be a square, as shown in Figure 13, or it may be another shape such as a rectangle.
[0065] Figure 14 is a schematic diagram illustrating the operation of the CABAC entropy encoder.
[0066] The CABAC encoder operates on binary data, that is, data represented by only two symbols, 0 and 1. Based on the already encoded data, the encoder performs a so-called context modeling process, which selects a "context," or probabilistic model, for the next data. Context selection is performed in a deterministic manner, based on already decoded data, without requiring any additional data (to identify the context) to be added to the encoded data stream passed to the decoder, so that the same decision is made in the decoder.
[0067] Referring to Figure 14, the input data to be encoded may be passed to the binary converter 1400 unless it is already in binary format. If the data is already in binary format, the converter 1400 is bypassed (by switch 1410 in the figure). In this embodiment, the conversion to binary format is actually performed by representing the quantized DCT coefficient data as a series of binary "maps". Binary maps will be described later.
[0068] The binary data may then be processed by one of two processing paths, a "normal" path and a "bypass" path (although schematically shown as separate paths, in some embodiments of the present invention described later, these can actually be executed in the same processing stage using only slightly different parameters). The bypass path uses a so-called bypass coder 1420 that does not necessarily utilize the same form of context modeling as the normal path. In some examples of CABAC coding, a bypass route can be selected when a series of data needs to be processed particularly quickly. However, in this embodiment, we will refer to two characteristics of so-called "bypass" data. The first characteristic is that bypass data is processed by the CABAC encoder (950,1460) using only a fixed context model that represents a 50% probability. The second characteristic is that bypass data pertains to a specific category of data. A specific example of such data is coefficient code data. If no bypass path is selected, the normal path is selected by switches 1430 and 1440, shown in the diagram, which operate under the control of the control circuit 1435. This includes data that is processed by the context modeler 1450 and subsequently by the encoding engine 1460.
[0069] The entropy encoder shown in Figure 14 encodes a block of data (i.e., data corresponding to a block of coefficients related to a block of residual images) as a single value if the entire block consists of zero-value data. For each block that does not fall into this category, i.e., a block containing at least some non-zero data, a "significance map" is created. The significance map indicates whether the corresponding coefficient within the data block is non-zero for each position in the data block being encoded. The importance map data itself, which is in binary format, is CABAC encoded. Using importance maps helps with compression because it eliminates the need to encode data for coefficients of magnitude indicated as zero by the importance map. The importance map can also include a special code indicating the last non-zero coefficient in a block. This allows all last high-frequency / trailing zero coefficients to be omitted from encoding. In the encoded bitstream, the importance map is followed by data defining the values of the non-zero coefficients specified by the importance map.
[0070] Furthermore, other levels of map data are created and CABAC encoded. One example is a map that defines, as a binary value (1=yes, 0=no), whether the coefficient data at a map location indicated as "non-zero" by the importance map actually has a value of "1". Another map specifies whether the coefficient data at a map location indicated as "non-zero" by the importance map actually has a value of "2". Yet another map indicates whether the data at these map locations, indicated as "non-zero" by the importance map, has a value of "3 or greater". And yet another map indicates the sign of the data value for data identified as "non-zero" (using predetermined binary notation such as 1 for +, 0 for -, or vice versa).
[0071] In embodiments of the present invention, significance maps and other maps are assigned in a predetermined manner to either a CABAC encoder or a bypass encoder, and all represent different attributes or value ranges of the same initial data item. In one example, at least the significance map is CABAC encoded, and at least a portion of the remaining map (such as the coded data) is bypass encoded. Thus, each data item is divided into respective subsets of the data, and each subset is encoded by a first (e.g., CABAC) and a second (e.g., bypass) encoding system. The nature of the data, CABAC, and bypass coding is such that for a given amount of CABAC-coded data, a variable amount of bypass data greater than or equal to zero is generated for the same initial data item. Therefore, for example, if the quantized, rearranged DCT data contains substantially all zero values, no bypass data may be generated, or only a very small amount of bypass data may be generated. This is because bypass data only relates to map locations where the significance map value is not zero. In another example, a considerable amount of bypass data may be generated in quantized reordered DCT data with many high-value coefficients.
[0072] In embodiments of the present invention, significance maps and other maps are generated from quantized DCT coefficients, for example, by the scanning unit 360, and subjected to a zigzag scanning process (or a scanning process selected from zigzag, horizontal raster, and vertical raster scanning) before being subjected to CABAC coding.
[0073] Generally, CABAC coding involves predicting the context of the next bit to be coded, i.e., a probabilistic model, based on other data coded previously. If the next bit is the same as the bit identified as "most likely" by the probabilistic model, coding the information that "the next bit matches the probabilistic model" can be coded with great efficiency. Since encoding "the next bit does not match the probabilistic model" is inefficient, deriving context data is crucial for good encoder operation. The term "adapt" means that the context or probabilistic model is adapted or changed during encoding in an attempt to provide a good match with the next data (which has not yet been encoded).
[0074] By simple analogy, the letter "U" is relatively rare in written English. However, it is very common in the position immediately following the letter "Q". Therefore, the probability model can set the probability of "U" to a very low value, but if the current letter is "Q", the probability model for "U" as the next letter can be set to a very high value.
[0075] CABAC coding, in the current configuration, is used for at least the significance map and the map indicating whether non-zero values are 1 or 2. Bypass processing—identical to CABAC coding in these embodiments, but used for at least the coded data and a map indicating whether a value is >2, given that the probability model is fixed with equal (0.5:0.5) probability distributions for 1s and 0s. For those data locations identified as >2, the actual value of the data can be coded using a separate so-called escape data encoding. This may include the Golomb-Rice coding technique.
[0076] CABAC context modeling and encoding processes are described in more detail in WD4: Working Draft 4 of High Efficiency Video Coding, JCTVC-F803_d5, Draft ISO I / EC 23008-HEVC; 201x(E) 2011-10-28.
[0077] Referring here to Figures 15 and 16, the entropy encoder, which forms part of the video encoding device, includes a first encoding system (an arithmetic encoding system such as the CABAC encoder 1500) and a second encoding system (such as the bypass encoder 1510), which are configured such that a particular data word or value is encoded into the final output data stream by either the CABAC encoder or the bypass encoder, but not by both.
[0078] In embodiments of the present invention, the data values passed to the CABAC encoder and the bypass encoder are respective subsets of ordered data values that are divided or derived from the initial input data (in this example, reordered quantized DCT data), representing different sets of “maps” generated from the input data.
[0079] The schematic diagram treats the CABAC encoder and bypass encoder as separate configurations. While this is often true in practice, in another possibility schematically shown in Figure 16, a single CABAC encoder 1620 is used as both the CABAC encoder 1500 and the bypass encoder 1510 in Figure 15. Encoder 1620 operates under the control of encoder 1630 to operate in an adaptive context model when in CABAC encoder 1500 mode (as described above), and in a fixed 50% probability context model when in bypass encoder 1510 mode.
[0080] A third possibility is to combine these two, in that two substantially identical CABAC encoders can be operated in parallel (similar to the parallel configuration in Figure 15). The difference is that the CABAC encoder operating as the bypass encoder 1510 has its context model fixed to a 50% probability context model.
[0081] The outputs of the CABAC coding process and the bypass coding process can be stored (at least temporarily) in buffers 1540 and 1550, respectively. In Figure 16, the switch or demultiplexer 1660 operates under the control of the mode signal 1630, routing the CABAC coded data to buffer 1550 and bypassing the coded data to buffer 1540.
[0082] Figures 17 and 18 schematically show an example of an entropy decoder that forms part of a video decoding device. Referring to Figure 17, the respective buffers 1710 and 1700 pass data to the CABAC decoder 1730 and bypass decoder 1720, and are arranged so that specific encoded data words or values are decoded by either the CABAC decoder or the bypass decoder, but not both. The decoded data is then reordered by logic 1740 into the appropriate order for subsequent decoding stages.
[0083] The schematic diagram in Figure 17 treats the CABAC decoder and bypass decoder as separate configurations. While this is often true in practice, in another possibility schematically shown in Figure 18, a single CABAC decoder 1850 is used as both the CABAC decoder 1730 and the bypass decoder 1720 in Figure 17. Decoder 1850 operates under the control of decoder 1860 to operate in an adaptive context model when in CABAC decoder 1730 mode (as described above), and in a fixed 50% probability context model when in bypass encoder 1720 mode.
[0084] As mentioned above, the third possibility is that two substantially identical CABAC decoders can be operated in parallel (similar to the parallel configuration in Figure 17), the difference being that the CABAC decoder operating as bypass decoder 1720 has a context model fixed with a 50% probability context model.
[0085] In Figure 18, the switch or multiplexer 1870 operates under the control of the mode signal 1860 to route the CABAC encoded data from buffer 1700 or buffer 1710 to decoder 1850, as appropriate.
[0086] Figure 19 schematically illustrates picture 1900 and is used to show various picture segmentation schemes relevant to the following discussion.
[0087] One example of picture partitioning is in slices or "canonical slices." Each canonical slice is encapsulated in its own Network Abstraction Layer (NAL) unit. Predictions within a picture (e.g., intra-sample predictions, motion information predictions, coded mode predictions) and entropy coding dependencies across slice boundaries are not permitted. In other words, a canonical slice can be reconstructed independently of other canonical slices within the same picture. A so-called tile defines horizontal and vertical boundaries for dividing a picture into rows and columns of tiles. In the way that corresponds to canonical slicing, in-picture prediction dependency is not allowed beyond the tile boundaries, and there is no entropy decoding dependency. However, tiles are not constrained to be contained within individual NAL units.
[0088] Generally speaking, a slice may contain multiple tiles, a tile may contain multiple slices, or a picture may contain one or more slices.
[0089] The schematic example in Figure 19 shows four slices 1910, 1920, 1930, and 1940, with slice 1940 containing two tiles 1950 and 1960. However, as mentioned, this is merely an arbitrary schematic example.
[0090] In some configuration examples, there is a threshold for the number of bins (either EP or CABAC) that can be encoded into a slice or picture, according to the following formula.
[0091] BinCountsinNalUnits <= (4 / 3) * NumBitsInVclNalUnits + (RawMinCuBits*PicSizelnMinCbsY) / 32 (Formula 1)
[0092] The right side of the expression depends on the sum of two parts. These are a constant value (RawMinCuBits*PicSizelnMinCbsY) for a particular image region, related to the size of the slice or picture, and a dynamic value (NumBitsInVclNalUnits) which is the number of bits encoded in the output stream of the slice or picture. Note that the value 4 / 3 represents the number of bins per bit.
[0093] RawMinCuBits is the number of bits in the minimum size (usually 4*4) of raw CUs, and PicSizelnMinCbsY is the number of CUs in the minimum size of a slice or picture.
[0094] When this threshold is exceeded, a CABAC zero word (3 bytes with the value 00 00 03) is appended to the stream until the threshold is met. Each such zero word increments the dynamic value by 3.
[0095] This constraint can also be expressed as follows: N <= K1 * B + (K2 * CU) Here, N = number of binarized symbols in the output data section. K1 is a constant. B = Number of encoded bytes in the output data section. K2 is a variable that depends on the characteristics of the minimum size coding unit adopted by the image data encoding device. CU = The size of the picture, slice, or tile represented by the output data section, which is expressed in the minimum number of coded units.
[0096] In the previously proposed example, this threshold check is performed at the picture and slice levels.
[0097] Alternatively, the following formula with the same threshold can be used, where the difference lies in the reference to bytes rather than bits. BinCountsinNalUnits <= (32 / 3) * NumByteslnVclNalUnits + (RawMinCuBits * PicSizelnMinCbsY) / 32 (Formula 2)
[0098] The threshold in Equation 1 (or the equivalent threshold expressed in Equation 2) can be applied uniformly, that is, regardless of other encoding or decoding parameters, attributes, etc.
[0099] One example of the technical justification for using such thresholds is that the threshold itself indirectly defines the maximum level of processing performance required by the real-time decoder. Real-time decoding relies on decoding each frame in time to achieve a specific output frame rate (e.g., frames per second). The constraint on the processing performance of the real-time decoder depends not on the amount of data that is decoded as a result, but on the speed at which the binary data must be decoded. To enable relatively simple rate control on the encoder side, the technique used to control this situation is to limit the number of binary values per slice or picture (e.g., as an upper limit, threshold, or constraint). The maximum processing performance required by the CABAC decoder can then be considered as the product or function of this constraint, and the speed at which the slice or picture must be decoded.
[0100] Therefore, it is important to set thresholds (or upper limits or constraints) appropriately. If the threshold is too high, it can cause problems with the implementation or proper operation of the real-time decoder. Similarly, the use of incorrect or inappropriate thresholds can lead to excessive use of padding data (e.g., dummy or other data that simply occupies a portion of the encoded data stream in order to increase the size of the encoded output data portion), which reduces the efficiency of encoding.
[0101] As noted with reference to Figure 19, a picture or slice can be divided into multiple tiles. One reason for doing this is to allow the use of multiple simultaneous (parallel) decoders.
[0102] Under the previously proposed configuration, each tile does not necessarily satisfy the threshold calculation discussed earlier. For example, if tiles are used or decoded individually like pictures, or if different tiles (with different quantization parameters or different sources, etc.) are combined together, there may be no guarantee that the combined slice or picture will comply with the above specifications.
[0103] To address this issue, in exemplary embodiments, the CABAC threshold is applied to the end of each tile, rather than to the end of each slice or image individually. Therefore, the threshold is applied at the end of encoding whether it is a tile, slice, or picture. With this in mind, we can assume that if each tile in an image conforms to the threshold, the entire picture must also conform, so in the case of a picture divided into slices or tiles, there is no need to apply the threshold again at the end of encoding the picture.
[0104] The terms “tile” and “slice” refer to independently decodeable units and represent the names used as of the priority date of this application. In the case of subsequent or other variations of the names, the arrangement is applicable to other such independently decodeable units. Therefore, in the exemplary configuration, the output data portion may be an independently decodeable data unit. For example, the image portion (represented by the output data portion) may be an image, a slice, or a tile.
[0105] To apply the above formula, the dynamic value represents the number of bytes encoded in the tile's output stream, while the fixed value depends on the number of minimum-size encoding units (CUs) within the tile.
[0106] Figure 20 schematically shows the apparatus configured to perform this test. Referring to Figure 20, at input 2000, a CABAC / EP encoded stream is received from the encoder. The padding data detector 2010, referring to the completion of a slice or tile, such as the end of the encoded slice or tile, detects at a predetermined stage whether the threshold calculation described above is in compliance. The controller 2020 controls the padding data generator 2030 in response to detection by the detector 2010 to generate padding data 2040, such as the CABAC zero word described above, and adds this to the stream by the coupler 2050 to form the output stream 2060. Zero word generation can also be performed by continuously monitoring whether the padding data detector 2010 complies with the threshold once the zero word is added, and after complying with the threshold, sending a signal back to the detector 2010 to cause the controller 2020 to stop generating the zero word.
[0107] In other examples, the controller 2020 may also function as a predictor configured to generate a prediction of whether the constraints are met by the output data section in relation to the embodiment described herein, while the output data section is being generated.
[0108] The predetermined stage may be, for example, every n encoded binary values (where n is an integer greater than or equal to 1), but in this embodiment, the predetermined stage is the end of the encoded output data section.
[0109] Accordingly, Figure 20 discloses an example of a padding data detector 2010 configured to detect whether constraints are met by the current output data section at a predetermined stage of encoding the current output data section, and a padding data generator 2030 configured to generate and insert sufficient padding data into the current output data section so that the output data section containing the padding data to be inserted satisfies the constraints.
[0110] Referring to Figure 21, in some examples, the controller 2020 includes an attribute detector 2070 configured to detect coding attributes applicable to a given output data section, and a selector 2080 configured to select constraints from two or more candidate constraints 2082 for use with a given output data section, depending on the detected coding attributes.
[0111] The controller 2020 may also include a comparator 2090 for comparing a threshold derived from the currently selected constraints with a detection by the padding data detector 2010 in order to derive a control signal for controlling the operation of the padding data generator 2030.
[0112] The attributes detected by detector 2070 may be, for example, encoding attributes (e.g., encoding mode or profile, e.g., enabling dependent quantization, which will be discussed below), which are operating modes in which the selection of quantization parameters used to quantize the current data value depends at least in part on the characteristics of the previously encoded data value and are contained in or associated with the subsequently encoded data stream, or represented by flag data (roughly shown as 2072) that can later be detected by the decoder. For example, flag data may be included in header data, such as the header data of the output data section (e.g., slice header data). The detector 2070 does not need to generate or insert flag data itself. This aspect of the process is shown in Figure 21 for schematic purposes only, for the benefit of this description. Attributes can be considered applicable to a given output data section, even if (as in some examples) the attribute is applicable to other output data sections (e.g., preceding or succeeding ones).
[0113] Therefore, the apparatus of Figures 7 and 14, operating according to the techniques described with respect to Figures 19 and 20 (and those described below), provides an example of an image data encoding apparatus. This apparatus is The system comprises a first data encoder 1450, 1460 and a second data encoder 1420, each configured to generate output data bits representing a binarized symbol from a sequence of symbols representing image data. The first data encoder is configured to generate output data bits representing coded symbols at a variable ratio of data bits to the number of coded symbols. The second data encoder is configured to generate a fixed number of output data bits representing each encoded symbol. The entropy encoder is configured to generate an output data stream (for example, using controller 2020 and / or controller 1435). This output data stream is subject to constraints that define an upper limit on the number of binarized symbols represented by the byte size of any individual output data section, with respect to the byte size of the output data section, so that the entropy encoder is configured to provide padding data for each output data section that does not satisfy the constraints, in order to satisfy the constraints and increase the byte size of the output data section. The device is An attribute detector 2070 configured to detect an encoding attribute applicable to a given output data section, The system includes a selector 2080 configured to select from two or more candidate constraints for use with a given output data section, depending on the detected encoding attribute.
[0114] Using the technique shown in Figure 14, the first data encoder / decoder may be a context-adaptive binary arithmetic coding (CABAC) encoder / decoder. The second data encoder / decoder may be a bypass encoder / decoder. The second data encoder / decoder may be a binary arithmetic coder / decoder using a fixed 50% probability context model.
[0115] Examples of applying thresholds or constraints are as follows:
[0116] Generally, the test is "Does the amount of generated output data meet the threshold test?" This can be performed, for example, at the end of encoding the output data portion.
[0117] The apparatus shown in Figure 7, operating according to the techniques discussed here, provides an example of an image data encoding device. The image data encoding device comprises an entropy encoder configured to selectively encode data items representing image data so as to generate encoded binarized symbols for consecutive output data sections. The entropy encoder is configured to generate an output data stream, which is subject to constraints that define an upper limit on the number of binarized symbols represented by the byte size of any individual output data section relative to the byte size of the output data section. To satisfy this constraint, the entropy encoder is configured to provide padding data to each output data section that does not satisfy this constraint, thereby increasing the byte-by-byte size of the output data section. The image data encoding device is An attribute detector configured to detect an encoding attribute applicable to a given output data section, The system includes a selector configured to select from two or more candidate constraints for use with a given output data section, depending on the detected encoding attribute.
[0118] Furthermore, this disclosure provides suitable decoding apparatus for decoding data signals generated by the methods or apparatus described herein.
[0119] Here are some further examples of constraints.
[0120] In the example above, a single threshold derivation or expression is consistently used. In the alternative example described below, a choice is implemented between two or more candidate thresholds or constraints. The choice may respond to one or more coding attributes (e.g., parameters, attributes, or modes, such as a flag, that are signaled from the encoder side to the decoder side, or parameters, attributes, or modes that can be derived in a corresponding or matching manner on both the encoder and decoder sides).
[0121] An example of a further candidate expression for the threshold or constraint is as follows: BinCountsinNalUnits <= 10 * NumByteslnVclNalUnits + (RawMinCuBits * PicSizelnMinCbsY) / 16 (Formula 3)
[0122] Therefore, although Equation 3 can be used uniformly as described above, in the embodiment, the selection is performed between two or more candidate expressions, e.g., Equations 1 / 2 and Equation 3. In some examples, the selection may depend on whether a predetermined encoded attribute that is signaled in or with the encoded data stream is in a first or second state on the encoder and decoder sides, each having a predetermined attribute state corresponding to the selection (e.g., Equations 1 / 2 or Equation 3).
[0123] An example of such an attribute is the so-called "dep_quant_enabled_flag." This is a slice header that indicates whether a technique called dependent quantization is enabled for the slice to which the slice header applies.
[0124] Instead of using `dep_quant_enabled_flag`, the actual availability of the tool depends on its availability rather than whether the `dep_quant` tool is enabled or not. Therefore, for example, in the case of profiles, the profile itself might define that `dep_quant_enabled_flag` should be off (not enabled). Alternatively, there might be no such constraint, allowing the `dep_quant` tool to be turned on (enabled) or off (disabled). Therefore, this choice can be made using profile constraints, rather than relying on whether the dep_quant tool is currently enabled or disabled.
[0125] Dependent quantization is defined in "Versatile Video Coding (Draft 5), Bross et al, JVET-N1001-v10, July 2019 (incorporated herein by reference)." See, for example, Section 8.7.3. This concerns the technique by which the decoding process selects between several possible quantization parameters or a set of quantization parameters in response to, for example, the properties of the previously encoded and decoded sample values (e.g., the parity property). Thus, when dep_quant_enabled_flag = 1 (enabled), such an ongoing dependent quantization selection takes place. When dep_quant_enabled_flag = 0 (disabled), such an ongoing dependent quantization selection does not take place. As mentioned earlier, the flag dep_quant_enabled_flag (as an example of a coding attribute) is provided, for example, in the slice header, so that enabling or disabling dependent quantization applies to the entire slice.
[0126] Different constraints may be involved when dependent quantization is used, or at least when it is possible, due to the empirical observation that applying dependent quantization can alter the expected relationship between the encoded binary and the decoded binary. Different constraints, such as those in equation (3), may be more appropriate for use with dependent quantization.
[0127] When multiple candidate constraints are applicable, the decoder is expected to be subject to design constraints in order to provide sufficient processing power, speed, or capacity to handle encoded data generated under the more difficult (or most difficult) of the different available constraints.
[0128] However, more generally, such arbitrary attributes (e.g., flags or parameters) can be used, whether explicitly within the data stream or whether they are signaled along with the data stream. For example, for each different instance of a so-called "profile," a different candidate expression can be selected, where the profile in this context defines a set or basket of parameters such as bit depth, chrominance sampling (e.g., 4:0:0 (monochrome), 4:2:0, 4:2:2, 4:4:4, etc.), and encoding type restrictions (e.g., in-image encoding only).
[0129] In exemplary embodiments, the relevant threshold derivation can be applied to the output data section, so attributes that define some aspect of the output data section can be conveniently used. Examples of output data sections in this context may include output data sections representing respective image parts such as slices, tiles, or pictures. Other examples of appropriate attributes include, in the case of slices, that the slice type is either intra-only or unrestricted (can include intra). As an example of tiles, consider the case of a composition of images from multiple sources (one tile per source). Each tile may have its own threshold derivation depending on the original encoding method. Alternatively, the attributes may be determined using the attribute values of the previous tile / output data section at the same position in the picture.
[0130] In some examples, as shown above, the constraint or threshold is defined by the following expression: N <= K1 * B + (K2 * CU) Here, N = Number of binarized symbols in the output data section, K1 is a constant. B = Number of encoded bytes in the output data section. K2 is a variable that depends on the characteristics of the minimum size coding unit adopted by the image data encoding device. CU = The size of the picture, slice, or tile represented by the output data section, which is expressed in the minimum number of encoded units.
[0131] It should be noted that this is, in fact, a generalization of Equations 1, 2, and 3 described above. An example of a list of candidate constraints that refer to K1 and K2 is as follows: [Table 1]
[0132] Regarding the example in Equation 5, the variable vcIByteScaleFactor can be expressed as follows: vcIByteScaleFactor = (32 + 4 * general_tier_flag) / 3 Here, general_tier_flag is an indicator of the encoding layer, which (in at least in some examples) changes as a flag value of 0 or 1. 0 indicates the so-called main layer, and 1 indicates the so-called high layer. For a given encoding level (representing the maximum dimensions of the image to be encoded), higher layers generally correspond to higher bitrate representations than the main layer. Thus, in this example, the image data encoding device is configured to operate with an encoding layer selected from at least two candidate encoding layers and to generate layer parameters that define the currently selected encoding layer (e.g., encoded in or at least in relation to the encoded image data or bitstream), where at least a constant K1 depends on the layer parameters. For example, a higher tier parameter may result in higher quality encoded output for a given image size, and the parameter K1 may increase depending on the tier parameter.
[0133] This means that the processes, circuits, code, or logic used to generate the thresholds are conveniently identical or substantially identical in each case, and the parameters K1 and K2 are simply modified with respect to each of the exemplary expressions (1 / 2) to (7). However, it is understood that one or more different expressions or expressions (or potentially different fixed thresholds) can be used, as between the different candidate thresholds or constraints described above.
[0134] Therefore, the candidate constraints 2082 can be stored or represented as pairs of (K1, K2) to be selected by the selector 2080. The pairs of (K1, K2) can be predetermined and stored, and the encoder and decoder, or the pairs (or instructions for a larger set of subsets of a given pair), can be transmitted from the encoder to the decoder as part of (e.g.) profile or parameter setting data. Next, the equation N <= K1*B + (K2*CU) can be tested with comparator 2090 using a commonly selected (K1, K2).
[0135] Figures 22A and 22B provide two schematic examples of the output data section, extending across the page and including the binary data 2200 and padding data 2210 provided and inserted by the technique described above. The amount of padding data can be determined by the specific image data being encoded (in terms of how efficiently it can be encoded) and the constraints in use.
[0136] Figure 23 is a schematic flowchart illustrating the image data encoding method, and this method is: (In step 2300) a step of selectively encoding data items representing image data to generate encoded binarized symbols for consecutive output data sections, (Step 2310) A step of generating an output data stream that is subject to constraints defining an upper limit on the number of binarized symbols represented by any individual output data section with respect to the byte-sized output data section, (In step 2320) In order to satisfy this constraint, the step of providing padding data to each output data section that does not satisfy this constraint in order to increase the byte-unit size of that output data section, (In step 2330) the step of detecting an encoding attribute applicable to a given output data section, (In step 2340) in response to the encoding attribute detected, a step of selecting a constraint from two or more candidate constraints for use with a given output data section and Includes.
[0137] Furthermore, exemplary embodiments provide an image decoder comprising a circuit configured to interpret an encoded signal generated by controlling an image data encoding device of any one or more embodiments described herein, and to output a decoded video image.
[0138] Furthermore, exemplary embodiments provide an image decoder that includes a circuit configured to interpret encoded signals generated by controlling first and second data encoders and one or more controllers of any of the embodiments described herein, and to output a decoded video image.
[0139] The embodiments described herein may be implemented in any suitable form, including hardware, software, firmware, or any combination thereof. The embodiments described herein may optionally be implemented at least partially as computer software running on one or more data processors and / or digital signal processors. Components and configuration requirements in any embodiment may be implemented physically, functionally, and logically in any suitable manner. In fact, functionality may be implemented in a single unit, in multiple units, or as part of other functional units. Accordingly, embodiments of the present disclosure may be implemented in a single unit or may be physically and functionally distributed among different units, circuits, and / or processors.
[0140] In light of the above teachings, it will be apparent that numerous modifications and variations of this disclosure are possible. Accordingly, it should be understood that, within the scope of the appended provisions, the technology can be implemented in ways other than those specifically described herein.
[0141] Each aspect and feature is defined by the following numbered clauses. (1) An entropy encoder configured to selectively encode data items representing image data so as to generate encoded binarized symbols of consecutive output data sections, An attribute detector configured to detect an encoding attribute applicable to a given output data section, A selector configured to select a constraint from two or more constraint candidates for use with the given output data section, depending on the detected encoding attribute. It is equipped with, The entropy encoder is configured to generate an output data stream, and this output data stream is subject to constraints that define an upper limit on the number of binarized symbols represented by the byte size of any individual output data section relative to the byte size of the output data section. To satisfy this constraint, the entropy encoder is configured to provide padding data to each output data section that does not satisfy the constraint, in order to increase the byte size of the output data section. Image data encoding device. (2) The entropy encoder is configured to selectively encode data items that are to be encoded by a first context-adaptive binary arithmetic coding (CABAC) coding system or a second bypass coding system in order to generate encoded binarized symbols. (1) The image data encoding device described above. (3) The image data represents one or more pictures, and each picture is (i) One or more slices within each network abstraction layer (NAL), (ii) Define the horizontal and vertical boundaries of each picture area, and zero or more tiles that are not constrained to be encapsulated within each NAL unit, It contains data representing, Each slice of the picture can be decoded independently of other slices of the same picture, and each tile can be decoded independently of other tiles of the same picture. The output data section includes one or more pictures, slices, and tiles. (1) The image data encoding device described above. (4) The second coding system is a binary arithmetic coding system that uses a fixed 50% probability context model. (1) or (2) the image data encoding device. (5) The above restrictions are, N <= K1 * B + (K2 * CU) (constraint equation 1) Here, N = the number of binarized symbols in the output data section. K1 is a constant. B = number of encoded bytes in the output data section, K2 is a variable that depends on the characteristics of the minimum size coding unit adopted by the image data coding device, CU = the size of the picture, slice, or tile represented in the output data section, which is expressed as the minimum number of coding units. Defined as An image data encoding device as described in any one of (1) to (4). (6) At least two candidate constraints are defined by constraint equation 1, Each of the series (K1, K2) is associated with each of the at least two candidate constraints, The selector is configured to select a series of (K1, K2) for the given output data section. (5) The image data encoding device described above. (7) The controller is configured to encode a representation of the encoding attribute applicable to the predetermined output data portion in relation to an output data stream representing the predetermined output data portion. An image data encoding device as described in any one of (1) to (6). (8) The image data encoding device includes a quantization unit configured to operate selectively in dependent quantization mode, The encoding attribute indicates whether the dependent quantization mode is valid or invalid with respect to a given output data portion. (7) The image data encoding device described above. (9) The entropy encoder is A detector configured to detect whether the constraints are satisfied by the current output data section at a predetermined stage in the encoding of the current output data section, A padding data generator configured to generate and insert sufficient padding data into the current output data section so that the output data section containing the padding data to be inserted satisfies the constraints, and including An image data encoding device as described in any one of (1) to (8). (10) The predetermined step is the end of the current output data section that has been encoded. (6) The image data encoding device described above. (11) A video storage device, capture device, and transmitting / receiving device comprising an image data encoding device as described in any one of (1) to (10). (12) A step of selectively encoding data items representing image data to generate encoded binarized symbols of consecutive output data sections, The steps include generating an output data stream that is subject to constraints defining an upper limit on the number of binarized symbols represented by any individual output data section with respect to the byte-sized output data section, To satisfy this constraint, the step of providing padding data to each output data section that does not satisfy this constraint, in order to increase the byte size of the output data section, A step of detecting an encoding attribute applicable to a given output data section, The steps include: selecting a constraint from two or more candidate constraints to use with a given output data section in response to the detected coding attribute; including Image data encoding method. (13) Computer software that, when executed by a computer, causes the computer to perform the method described in (12). (14) A machine-readable non-temporary storage medium for storing the computer software described in (13). (15) A data signal containing encoded data generated according to the method described in (12). (16) An image data decoder configured to decode the data signals described in (15).
Claims
1. A buffer stores an input data stream representing image data subject to a constraint that defines an upper limit on the number of binarized symbols that can be represented by each output data unit, according to the byte size of the output data unit, wherein padding data is provided for input data units that do not satisfy the constraint, thereby increasing the byte size of the input data unit to satisfy the constraint, and this constraint is selected based on a mode or profile of the input data stream that satisfies the design constraints of the entropy decoder. Decode the data items of the input data stream representing the image data using a first context-adaptive binary arithmetic coding (CABAC) decoding system or a second bypass decoding system, and generate decoded binarized symbols by determining the ability to handle selected constraints using the attributes in the input data stream. and Image data encoding method.
2. The aforementioned constraints N <= K1 * B + (K2 * CU) (Constraint equation 1) Here, N = the number of binarized symbols in the output data section. K1 is a constant. B = number of encoded bytes in the output data section, K2 is a variable that depends on the characteristics of the minimum size coding unit adopted by the image data coding device, CU = the size of the picture, slice, or tile represented in the output data section, which is expressed by the minimum number of coding units. Defined as The image data encoding method according to claim 1.
3. At least two candidate constraints are defined by constraint equation 1, Each of the series (K1, K2) is associated with each of the at least two candidate constraints, The selector selects a series of (K1, K2) for the given output data section. The image data encoding method according to claim 2.
4. The mode or profile is indicated by a flag in the data stream. The image data encoding method according to claim 1.
5. The attributes in the data stream are flags derived by the decoding circuit. The image data encoding method according to claim 1.
6. A flag that determines the selection of constraints is used to determine the ability of the decoder circuit to deal with the selected constraints. The image data encoding method according to claim 4.
7. The flags that indicate profiling define a set of restrictions that apply to the input stream. The image data encoding method according to claim 4.
8. A computer program that, when executed by a computer, causes the computer to perform the method described in Claim 1.
9. A computer-readable storage medium for storing the computer program described in Claim 8.
10. A buffer configured to receive and store an input data stream representing image data subject to constraints defining an upper limit on the number of binarized symbols that can be represented by each output data unit with respect to the byte size of the output data unit, wherein padding data is provided for each input data unit that does not satisfy the constraints, and the byte size of the input data unit is configured to increase to satisfy constraints selected based on the mode or profile of the input data stream that satisfy the design constraints of the entropy decoder, A decoding circuit generates a decoded binarized symbol by decoding the data items of the input data stream representing the image data using a first context-adaptive binary arithmetic coding (CABAC) decoding system or a second bypass decoding system, and by determining the ability to deal with selected constraints using the attributes in the input data stream. An image data encoding device having the following features.
11. The aforementioned constraints N <= K1 * B + (K2 * CU) (Constraint equation 1) Here, N = the number of binarized symbols in the output data section. K1 is a constant. B = number of encoded bytes in the output data section, K2 depends on the characteristics of the minimum size coding unit adopted by the image data encoding device. It is a variable that does, CU = the size of the picture, slice, or tile represented in the output data section, which is expressed by the minimum number of coding units. Defined as The image data encoding device according to claim 10.
12. At least two candidate constraints are defined by constraint equation 1, Each of the series (K1, K2) is associated with each of the at least two candidate constraints, The selector is configured to select a series of (K1, K2) for the given output data section. The image data encoding device according to claim 11.
13. The aforementioned mode or profile is indicated by a flag within the data stream. The image data encoding device according to claim 10.
14. The attributes within the data stream are flags derived by the decoding circuit. The image data encoding device according to claim 10.
15. The flags that determine the selection of constraints are used to determine the decoder circuit's ability to handle the selected constraints. The image data encoding device according to claim 14.
16. The flags that indicate profiling define a set of restrictions that apply to the input stream. The image data encoding device according to claim 14.
17. A video storage device, capture device, and transmitting / receiving device including the image data encoding device according to claim 10.