Method and apparatus for complexity control in high throughput JPEG2000 (HTJ2K) encoding

By using statistical analysis and prediction methods to determine quantization parameters for each code block, the proposed method addresses the challenge of managing coding complexity in image and video coding, achieving efficient and high-quality encoding.

JP7691752B2Active Publication Date: 2025-06-12KAKADU R & D PTY LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022523830
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-10-25
Filing Date
2020-10-26
Publication Date
2025-06-12
Estimated Expiration
2040-10-26

AI Technical Summary

Technical Problem

Conventional image and video coding algorithms face challenges in managing coding complexity, particularly in achieving target compression sizes without high computational and memory costs, while maintaining image quality.

Method used

The proposed method involves collecting local or global statistics for each subband, performing spatial transformation, and generating predictions of subband sample statistics to determine a single quantization parameter (QP) for each code block, which allows for the derivation of the coarsest bitplane to be generated, thereby optimizing complexity constraints.

Benefits of technology

This approach enables efficient complexity-constrained encoding of JPEG2000 code streams, achieving significant improvements in image quality and throughput while maintaining low computational complexity, even in low-memory configurations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007691752000071
    Figure 0007691752000071
  • Figure 0007691752000072
    Figure 0007691752000072
  • Figure 0007691752000073
    Figure 0007691752000073
Patent Text Reader

Abstract

A method for managing encoding complexity for image and video coding, e.g. for algorithms belonging to the JPEG2000 family of standards, where the encoding process targets a given compressed size (i.e. total encoded length) of each frame of an image or video sequence. A set of methods for complexity-constrained coding of HTJ2K codestreams is described, which involves collecting local or global statistics for each subband (rather than each codeblock), generating predictions of the statistics for subband samples not yet produced by the spatial transform and quantization process, and using this information to generate a global quantization parameter, which can derive the coarsest bitplane to be generated for each codeblock. The coding length estimates are generated in a way that latency and memory are separately optimized for coded image quality, while maintaining low computational complexity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image and video coding. More particularly, without limitation, the present invention relates to the management of coding complexity, and in particular, but not exclusively, to the management of the coding complexity of algorithms belonging to the JPEG2000 standard family, where the coding process targets a given compression size (i.e., the total length of the code) for each frame of an image or video sequence.

[0002] The embodiments described in this disclosure are particularly beneficial when applied to the coding technique known as High Throughput JPEG2000 (HTJ2K) described in Part 15 of the JPEG2000 standard family, formally ITU-T Rec T.814|ISO / IEC15444-15. HTJ2K defines a new "HT block coding algorithm" that can be used in the compression techniques of other parts of the JPEG2000 standard family. However, the embodiments described in this disclosure may also have advantages when coding video content with the original block coding algorithm of JPEG2000 Part 1.

Background Art

[0003] In most conventional compression algorithms, the only way to achieve the target compression size is to adjust a set of quantization parameters, usually via a master parameter such as a quality factor (JPEG) or QP parameter (H.264 / AVC, H.265 / HEVC). This is done either by performing iterative coding at the entire image level until the compression size approaches the target value or by performing online progressive adaptation such that the compression quality varies across the entire image. In the first case, the computation and memory consumption can be very high, and in the second case, the image quality may degrade and the compression size cannot be definitely constrained.

[0004] In contrast, JPEG2000 can achieve the target compression size without the need for iteration or adaptation of quantization parameters. This is because each code block in each subband has an embedded representation, and after the encoding is completed, by simply truncating the embedded bitstream of each block, the compression size and distortion can be exchanged in a nearly optimal form. This is usually achieved using the Post Compression Rate-Distortion optimization (PCRD-opt) algorithm, which is described in conjunction with the original Embedded Block Coding with Optimal Truncation (EBCOT) algorithm that is the basis of JPEG2000.

[0005] Recently, a new Part 15 has been added to the JPEG2000 standard family. Part 15, also known as "High Throughput JPEG2000", describes a new high throughput block coding algorithm. For convenience, here the original embedded block coding algorithm is called "J2K-1", and the new algorithm is called High Throughput "HT". Different from J2K-1, the HT algorithm does not generate a complete embedded bitstream for each code block. However, the HT algorithm generates a set of partial embedded coding paths that are organized into a so-called "HT set". A single HT set includes an HT Cleanup coding path, an HT SigProp coding path, and an HT MagRef coding path, which can be directly associated with the Cleanup, SigProp, and MagRef coding paths generated by the J2K-1 block coder.

[0006] The relationship between the J2K-1 encoding path and the HT encoding path is shown in Figure 1. Each HT set is associated with a base bitplane index p. The HT Cleanup path of that set encodes all samples in the code block to a precision related to the size of bitplane p, while the HT SigProp and HT MagRef encoding paths, if present, refine the precision of specific samples to the next finer bitplane p-1. Thus, these last two paths are known as HT refinement paths. J2K-1 does the same except that if the HT Cleanup path does not fully encode all samples to bitplane p (not embedded), the corresponding J2K-1 Cleanup path refines all samples to the precision of bitplane p taking into account all the information provided by the previous encoding path (embedded).

[0007] The advantage of the HT block encoding algorithm is that it can be run with a much higher throughput in both software and hardware and consumes much less computational energy. In decoding, it is only necessary to decode one HT set. Even when multiple HT sets are encoded, typically only one of them is included in the final code stream and it is always sufficient for the decoder to process at most one HT set per code block, i.e., one HT Cleanup path (if present) within the same HT set, as well as any HT SigProp and HT MagRef refinement paths.

[0008] The HT block coding algorithm of the encoder provides a broader opportunity to optimize the trade-off between complexity / throughput and image quality. Figure 2 is a diagram showing the elements of the HTJ2K encoder. HTJ2K substantially preserves the existing architecture and code stream syntax of JPEG2000. The image first undergoes any necessary multi-component transformation and / or non-linear point transformation as enabled by Part 1 or Part 2 of JPEG2000, and then the transformed image components are processed by a reversible or irreversible Discrete Wavelet Transform (DWT), which decomposes each component into a hierarchy of detail sub-bands and one base (LL) sub-band.

[0009] All sub-bands are divided into blocks of size 4096 samples or less, with typical dimensions of 64×64 or 32×32, and very wide and short blocks such as 1024×4 are also important for low-latency applications. Each block is quantized individually (if irreversible) and encoded to produce a block bit stream containing one or more encoded passes.

[0010] In the encoder, an optional Post Compression Rate-Distortion optimization (PCRD-opt) phase is used to discard the generated encoded passes to achieve rate or distortion targets that can be global (across the entire code stream) or local (a small window of code blocks). Finally, the bits belonging to the selected encoded passes from each code block are assembled into J2K packets to form the final code stream.

[0011] In both the J2K-1 and HT block coders, the encoder can drop any number of trailing coding passes from the information included in the final code stream. In fact, the encoder need not generate such coding passes in the first place if it can reasonably predict that such coding passes will be dropped. The strategy for doing this is described in "Software architectures for JPEG2000" by D. Taubman, Proceedings of the IEEE International Conference on DSP, Santorini, Greece (2002), and is routinely deployed at least in software implementations.

[0012] In the HT block coder, as long as the first coding pass output is a Cleanup pass, both the leading and trailing coding passes can be dropped (or not generated) by the encoder. As a result, the HT encoder typically only needs to generate just six coding passes corresponding to two consecutive HT sets such as those identified as HT set-1 and HT set-2 in FIG. 1. Subsequently, the PCRD-opt stage shown in FIG. 2 selects up to three passes of the coding passes generated from each code block for inclusion in the final code stream, and the selected passes belong to a single HT set.

[0013] In some cases, the encoder need not generate multiple HT Cleanup passes. This is certainly true for lossless compression where only the Cleanup pass for p = 0 is relevant, and this pass belongs to a degenerate HT set identified as the "HT Max" set in FIG. 1 and may not have refinement passes. During lossy compression, the distortion associated with the HT Max set can be set to the desired level of image quality in much the same way that quantization is used to control compression in JPEG and most other media coders, depending on the quantization parameter.

[0014] As described above, there are multiple ways for the HTJ2K encoder to compress an image or video source. The simplest approach is to generate only the single-pass HT Max set and manage the trade-off between image quality and compression size by modulating the quantization parameter.

[0015] Conversely, the encoder can generate all possible HT coding paths (one HT set per significant magnitude bitplane of each code block), leave the determination of the optimal point to discard the quality of each code block to the PCRD-opt rate control algorithm, and then select the Cleanup path, SigProp path, and MagRef path (at most one of each) that need to be included in the final code stream for the determined discard points of each code block. This is a complete waste of both computation and memory. In an optimized implementation, this approach is still several times computationally advantageous compared to the J2K-1 algorithm (e.g., 4 to 5 times faster), but because redundant information is included in multiple HT Cleanup paths, the cost of temporarily buffering the encoded data in memory is much higher than in the case of J2K-1. For reference, this is called "HT-Full" encoding.

[0016] The applicant's previous international application AU2019 / 051105 (currently published as International Publication No. 2020 / 073098) describes various methods for determining the number of leading coding paths to drop (i.e., the coarsest HT sets to generate) in a video coding application. In the first method, the encoder uses the information collected from the previous frame to establish constraints on the coding length of the generated HT coding paths for each code block within the current frame, and uses iterative coding techniques to ensure that at least some paths satisfy the constraints, after which the PCRD-opt algorithm is executed. This method has the difficulty that the number of coding paths that need to be generated for each code block cannot be deterministically limited in advance.

[0017] In the second method described in International Publication No. WO 2020 / 073098, various attributes of the PCRD-opt decisions made for each code block in the previous frame are recorded for use in determining an appropriate range of encoding paths to generate for the same code block in subsequent frames. As a result, the set of generated encoding paths adapt over time with higher or lower precision on a per-code-block basis. The purpose is to provide an appropriate range of options for the current frame's PCRD-opt algorithm while constraining the number of paths generated for any given code block in a deterministic manner. For reference in later experimental comparisons, this method is referred to herein as the "PCRD-Stats" method. The "PCRD-Stats" method has the drawback that it cannot respond quickly to changes in scene complexity over time, such as scene cuts. Both of these methods are only suitable for video encoding, as opposed to still image encoding.

[0018] International Publication No. WO 2020 / 073098 describes a further method that uses model-based techniques to convert the statistical values of the quantized subband statistics of each code block into estimates of the encoding length and distortion at each of a large set of truncation points. The estimated distortion-length characteristics of each code block are fed into a coarse PCRD-opt algorithm that estimates an approximately optimal truncation point for each code block based on the target compression total length. These estimated truncation points are then used to determine the range of encoding paths to actually generate, and the result is fed into the full PCRD-opt stage. This method is complex to implement and may require a large amount of memory to buffer the subband samples between the time when the statistical values are collected for the first (coarse) PCRD-opt stage and the time when the code block samples are actually encoded. SUMMARY OF THE INVENTION

[0019] Embodiments of the present invention describe a new series of methods for complexity-constrained coding of an HTJ2K code stream, including collection of local or global statistics for each subband (not each code block), spatial transformation, and generation of predictions of subband sample statistics not yet generated by the quantization process, and use of this information for global quantization parameter generation, from which the coarsest bitplane to be generated for each code block can be derived in a straightforward manner. In one embodiment, the application of the method is performed online (i.e., dynamically) when subband samples are generated to determine a set of HT sets to be generated for each code block for which quantized samples are available. The method can also be deferred until all subband samples for an image or video frame are generated and buffered in memory, at which point a set of HT sets to be generated for all code blocks in the image or frame, and thus the coarsest bitplane, are generated, after which the encoding itself can be performed. These variations support a variety of applications and deployment platforms, including low-memory and high-memory configurations. Embodiments describe prediction methods that can be used to achieve low-memory encoding to a target compression size even for still images, and efficient adaptive prediction methods that can utilize temporal information in video encoding applications while maintaining robustness to rapid changes in scene complexity are disclosed.

[0020] The differences between the embodiments described herein and the methods previously described for complexity-constrained coding using a partially embedded block coding algorithm include the following (see also "FBCOT: a fast block coding option for JPEG2000" by D. Taubman, A. Naman, and R. Mathew, SPIE Optics and Photonics: Applications of Digital Imaging, San Diego (2017)). 1. Embodiments of the novel method include determining, for each code block, a single quantization parameter (QP) that depends on the target compression size, along with a mapping that is independent of the target size, from the QP value to the coarsest bitplane. 2. Embodiments of the novel method include collecting a simple set of statistics for each subband, from which it is possible to estimate the compression size for each value of the QP value described above. 3. Embodiments of the novel method include predicting statistical values related to subband samples that have not yet been generated, such that the QP value can be periodically updated based on the estimated compression size derived together from the observed samples and the predicted samples. 4. In the case of video applications, information from previous frames in the sequence is incorporated into the method of the embodiment via an adaptive prediction of the statistics of the invisible subband samples in the current frame, and the prediction is formed using both spatial and temporal inference.

[0021] Embodiments described herein have applications in low-memory and high-memory software-based encoding arrangements, including GPU configurations, as well as in low-latency and high-latency hardware arrangements. The embodiments enable the separate optimization of latency and memory with respect to the encoded image quality while maintaining a low computational complexity. In experimental studies, embodiments of the present invention can significantly outperform previously reported encoding strategies with complexity constraints, both in terms of image quality and throughput / complexity. The main focus of the embodiments is on complexity-constrained HTJ2K encoding of images and videos, but the embodiments can also be used to improve the robustness of conventional (i.e., J2K-1) JPEG2000 encoding of videos.

[0022] The present invention provides a method for complexity-constrained encoding of JPEG2000 code streams, including JPEG2000 and high-throughput JPEG2000 code streams that are subject to a target full-length constraint. The method is as follows: a. Collecting information regarding sub-band samples generated by spatial transformation; b. Generating an estimated value of the coding length from the collected information for a plurality of potential bit-plane truncation points; c. After mapping quantization parameters (QP values) to the base bit-plane indices of each associated code block of each sub-band, determining QP values from these length estimates such that the overall coding length estimated when truncating at these bit-plane indices has a good chance not to exceed the target total length constraint; d. Mapping the QP values to the base bit-plane indices of each code block; e. Encoding each associated code block with an accuracy associated with the corresponding base bit-plane index and encoding one or more additional coding passes from each code block; f. Subjecting all the coding passes thus generated to a post-compression rate-distortion optimization process to determine a final set of coding passes output from each code block as a compression result, including.

[0023] One embodiment describes a method for generating an estimated value of the coding length. One embodiment describes a method for incrementally determining QP parameters using the prediction of the estimated coding length of unobserved sub-band samples to reduce the amount of memory required to buffer sub-band samples before block coding. One embodiment extends the prediction method to incorporate a robust combination of spatial and temporal prediction for video applications. One embodiment includes the application of the method disclosed above to low-latency image and video coding.

[0024] The present invention further provides an apparatus for coding a code stream with complexity constraints, the apparatus comprising an encoder configured to implement the above method.

[0025] The present invention further provides a computer program comprising instructions for controlling a computer to implement the above method.

[0026] The present invention further provides a non - volatile computer - readable medium storing the computer program as described above.

[0027] The present invention further provides a data signal comprising the computer program as described above.

Brief Description of the Drawings

[0028] The features and advantages of the present invention will become apparent from the following description of its embodiments, by way of example only, with reference to the accompanying drawings.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Embodiments for Carrying Out the Invention

[0029] In the following description of embodiments of the present invention, five "embodiments" are discussed. Embodiment 1: QP-based complexity control using an estimated value of the coding length

[0030] An important principle underlying this embodiment is that an almost optimal set of quantization parameters for the transformed image can all be described in terms of a single global parameter, herein identified as QP (note that this is not the same as the QP parameter used in modern video codecs such as H.264 / AVC, H.265 / AVC, or AV1, but serves a related role). First, this property will be explained.

[0031] Let x be the vector representing all samples in the image, and y[n] be the 2D sequence of subband samples from transform subband b (indexed by n ≡ [n 1 ,n 2 ). Then the relationship between the transform domain representation and the image domain representation can be expressed as follows. b [n].

Equation

Equation

Equation

Equation

[0032] More generally, it often aims to minimize a visually weighted distortion metric and can be expressed as follows.

Number

[0033] At high bit - rates, the general distortion - rate model in the sub - band sample quantization process is as follows.

Number

Number

Number

Number

Number

[0034] For the quantization step size selected for sub - band b, let Δ b be the number P of the least - significant bit - planes discarded from the quantized sub - band samples. b Then, it becomes as follows:

Number

Number

[0035] Thereafter, interpret the first and second terms on the right - hand side of the above equation as the sub - band - specific bias parameter β b and the global quantization parameter QP, respectively, as follows.

Number

[0036] The complexity control method of the present invention assigns the "base Cleanup path" indicated by Cup0 in the code block of sub - band b so as to correspond to the bit - plane.

Number

[0037] Since these preparations are complete, the core complexity control method of the present invention shown in FIG. 3 can be described.

[0038] In each code block of subband b, starting from bit plane P b at most Z coding paths are generated, where a preferred value of Z is 6. The value of P b is determined from Equation (1) using a QP value that is quantized to an integer multiple of G, where G can be interpreted as the number of "grid" points between consecutive integers. In FIG. 3, G = 4, which is a good choice. For G, QP can be expressed using an integer F as follows,

Number

Number

Number

[0039] To find the value of F (i.e., QP), the number of bytes associated with each candidate value p in P b is estimated, and an "estimated length record" (or vector) L p (b) including the estimated length L (b)is formed. The method for estimating the actual length, as will be described later in this specification, forms Embodiment 2 of the present invention. However, here, it is important to note that these estimated values are conservative, that is, the total number of encoded bytes related to the Cleanup path in the bit plane p across all code blocks of the sub-band b is close to L p (b) but not greater than it.

[0040] Each estimated length record L (b) is replicated G times and biased to form the following.

Number

Number

Number

Number

[0041] The extended length vector

Number

Number

Number

[0042] Obviously, when the integer F is selected to be equal to f, L f is the (conservative) estimated number of encoded bytes generated by the base Cleanup path Cup0 for all code blocks. The QP selection operation (the box shown by reference number 5 in Figure 3) simply selects the following.

Number

[0043] In the present invention, it is valuable to emphasize the fact that the length estimation process itself is not the basis for rate control (i.e., generating a code stream with a desired encoding length). The length estimation process ideally significantly underestimates the encoding length by a large margin, and the method of estimating the encoding length is preferably very simple and thus not suitable for reliable rate control. Instead, rate control is performed by a post-compression rate distortion optimization procedure that utilizes the availability of multiple truncation points for a code block, i.e., multiple encoding lengths and the associated distortion for a code block. In fact, one source of the idea of the present invention is the inventors' recent discovery (experimentally verified later in this specification) that it is possible to devise a low-complexity length estimator that has a very high probability of having the desired level of conservativeness.

[0044] The complexity control method described herein targets high-throughput JPEG2000 using the HT block coder, but can be applied to other media encoding systems. Most notably, using the same complexity control method, the number of encoding paths generated by the J2K-1 block encoding algorithm can be limited, such that for any given code block within subband b, at most Z-1 paths are generated beyond the Cleanup paths associated with base bitplane P b up to b . Since the encoding efficiency of the J2K-1 block coder is typically similar to that of the HT block decoder (e.g., about 10% better), the same method for estimating the encoding length can be used in both cases.

[0045] The main difference between the HT block encoding algorithm and the J2K-1 block encoding algorithm is that since J2K-1 is fully embedded, the J2K-1 block coder must generate all encoding paths associated with bitplanes coarser than base bitplane P b as long as the code blocks contain significant samples in those coarser bitplanes, whereas the HT block coder does not need to do so. Nevertheless, the fact that encoding can be stopped at Z-1 paths after the Cleanup path at bitplane P b can still represent a significant computational savings compared to simply generating all possible encoding paths for each code block.

[0046] This method can also be used to generate a code stream in which code blocks using the HT block encoding algorithm and code blocks using the J2K-1 block encoding algorithm are intermixed. Embodiment 2: Encoding Length Estimation Using Subband Sample Statistics

[0047] Here, as implemented by the boxes denoted by reference numerals 6 and 7 in FIG. 3, the estimated length L p (b)Pay attention to the problem of how to form it. This method models the encoding cost of an algorithm for encoding the following quantization indices.

Number

Number

[0048] The modeled algorithm encodes only the magnitude μ b,p [n], and the sign χ b [n] of the samples with non-zero magnitude. In a preferred embodiment, the modeled algorithm is a rough approximation of the actual encoding algorithm. In particular, the modeled algorithm of the HT Cleanup encoder intentionally omits many of the features that make the actual HT Cleanup algorithm efficient in order to obtain a conservative estimate of the encoding length L p (b) This specification describes specific embodiments that are computationally efficient and actually effective.

[0049] The first step of the estimation procedure involves collecting the quad significance statistic C b,p for each bitplane p. Specifically, the subband samples are such that the sample y b [n] belongs to the quad q, and q ≡ [q 1 , q 2 is indexed into 2×2 quads.

Number

Number

Number

[0050] Note that the statistic C b,p can be calculated without explicitly determining the individual μ b,p [n] values. It is sufficient to first form and quantize the maximum magnitude for each quad to obtain the following.

Number

[0051] Then, for each candidate bitplane p,

Number

Number

[0052] The second step of the estimation procedure is the statistic C b,pincludes converting to the estimated number of bytes. The HT Cleanup encoding algorithm generates three byte streams known as the MagSgn byte stream, the VLC byte stream, and the MEL byte stream. The number of bits packed from the quad q into the MagSgn byte stream depends on the precision boundary of the quantized magnitude within the quad. In the actual algorithm, this boundary is based on the so-called "magnitude exponent" E b,p [q], and E b,p [q] - 1 is the number of bits required to represent μ b,p [n] - 1 for any sample within the quad, and the additional 1 indicates that the sign bit of the non-zero sample needs to be transmitted. In the case of the simplified model here, it is convenient to use the quantities as follows,

Number

Number

[0053] In addition to the magnitude and sign bits, P b,p[q]Adopt a simple model for the cost of transmitting values. The actual HT Cleanup encoder, as a first reason, has a magnitude exponent E b,p [q]that is not the same as the P b,p [q]value (there is an offset of 1 for the bounded magnitude), and as a second reason, since the HT block coder has complex inter-sample (not just between quads) dependencies in the way it transmits the exponent boundaries, it differentially transmits magnitude exponent boundaries different from those for transmitting the P b,p [q]value. Although it is possible to attempt to model all of this, the preferred embodiment instead relies on the assumption that a simple coder that describes the same information should provide an upper bound on the coding length of the actual Cleanup algorithm. The simple coder recommended here has an adaptive run-length coding of quad significance symbols with a coding mechanism that is assumed to achieve the zero-order entropy of the quad significance symbol whenever the significance probability is less than 0.5, combined with (via runs of non-significant quads) a unary (comma) code for the P b,p [q]for each significant quad. Since the adaptive run-length coder cannot use less than 1 bit per run, when the significance likelihood is greater than 0.5, it degenerates to output a single significance bit for each quad. These two embodiments (quad significance coding and unary coding of the P b,p [q]of significant quads) are each somewhat similar to the information transmitted via the MEL and VLC bitstreams of the HT Cleanup coder.

[0054] In the case of the unary code, the first bit indicates whether P b,p [q]>1 conditioned on the quad being significant (i.e., P b,p [q]>0), and the second bit indicates whether P b,p [q]>2 conditioned on P b,p [q]>1, and so on. Thus, the total number of unary code bits for a subband is simply as follows.

Number

[0055] In the case of run-length encoding of importance, the number of bits is approximated as follows.

Number

[0056] In some embodiments, the adaptability of the run-length encoding procedure used by the HT Cleanup pass is to first accumulate the quad significance statistics over the individual line pairs j within the subband,

Number

Number

Number

[0057] In any case, the final estimated number of bytes for bitplane p of subband b is simply formed by adding the three components developed above and dividing by 8, resulting in the following equation.

Number

[0058] In practice, L b,p always overestimates the number of bytes required by the HT Cleanup encoding of bitplane p, but it has been observed that it tends to be smaller than the number of bytes required by the HT Cleanup encoding of the next finer bitplane p - 1. This is a property required for the overall complexity control algorithm of FIG. 3 to succeed when the number of encoding paths generated for each code block is at least Z = 4. For each subband b, the maximum bitplane P b max for which any quad is significant exists. For all p > P b max and C b,p = 0, the only non - zero contribution to equation (6) is a very small cost R b,p to convey the fact that all quads are not significant. In particular, for p > P b max the value of L b,p obtained using the above procedure depends at most on the subband size. P b max itself is data - dependent, but there is a clear boundary P b that depends only on the quantization step size Δ b max the properties of the transform used to generate the subband samples, and the bit depth of the original image sample values, such as P b bound ≤ P b bound . Embodiments of the present method can use this boundary to determine the number of quad significance statistics C b,p that need to be collected for each subband b.

[0059] In some applications, it may be desirable to impose a fixed limit S on the number of quad significance statistics C b collected for subband b, regardless of Δ b,p the transform characteristics or the image bit depth. This is for p ≪ Pb max By utilizing the fact that the following relationships tend to hold for (i.e., with very high precision), the effectiveness of the present method can be achieved without significantly degrading it.

Number

[0060] This is because, with very high precision, most samples within the sub - band become significant, C b,p ≈Q b , (V b,p - V b,p+1 )≈Q b , (M b,p - M b,p+1 )≈4Q b and (R b,p - R b,p+1 )≈0. Using this relationship, the embodiment can collect the statistical values C b,p and explicitly calculate the length estimate L b,p for only those bit - planes p as follows.

Number

[0061] The reader will understand that the above-described method for estimating the coding length is only one of many related methods that can be used to provide a conservative model of the number of bytes generated by the HT Cleanup coding procedure. More sophisticated models can be used that more accurately mimic the behavior of an actual encoder, but practical experience suggests that these may not be justified considering the fact that the HT Cleanup encoder itself has a low complexity and the above very simple model is sufficient even when the number of generated coded paths Z is small. Embodiment 3: Online QP Adaptation Using Prediction Statistics

[0062] In the complexity control method described in Embodiment 1, before the QP parameter can be calculated, it is necessary to estimate the coding length from the statistical values collected from all subband samples. Thereby, the coded path generated by the block coding process is determined. As a result, it is necessary to buffer in memory the entire image, its quantized subband samples, or some equivalent data set before the block coding process can be started. In many cases, even if the computational complexity is low, the memory complexity becomes high.

[0063] Embodiment 3 avoids high memory complexity by dynamically updating the QP value based on the subband samples actually generated by the spatial conversion process. FIG. 4 illustrates this method, identifying the role played by the other subbands while focusing only on one subband b. The spatial conversion (here a discrete wavelet transform) is pipelined so that the entire image does not need to be buffered in memory. As is well known, image lines can be incrementally pushed into a discrete wavelet transform (DWT) in a top-down fashion, in which case only a moderate amount of internal state memory is used to incrementally generate the lines of subband samples for each subband b. Of course, bottom-up and column-wise incremental pushes of the image data can also be realized similarly when appropriate, but in most applications the image data arrives in raster scan order, so the method in that case is described here in particular. The subband samples are collected in a memory buffer and consumed therefrom by the block encoding process.

[0064] It is useful to collect the subband samples in stripes, where each stripe represents the total number of code blocks of the subband (typically one row of code blocks). As shown in the figure, at any given point in the process, four categories of subband samples can be identified as follows. 1. The "active stripe" k corresponds to subband samples that have been generated by the conversion and are ready to have a QP value (equivalently, an F value) assigned, enabling the base Cleanup bitplane P b to be assigned and allowing the block encoding of these samples to proceed. 2. The "dispatched stripe" corresponds to previously active subband samples for which a QP value (i.e., an F value) has already been assigned, and thus the base Cleanup bitplane P bIt corresponds to these. These samples may already be encoded, but this is not a strict requirement. In a parallel processing environment, a code block can be distributed to a simultaneous processing engine that can span two or more stripes so that one or more dispatched stripes can still be executing while the QP value of the active stripe is being determined. The methods described herein do not require strict synchronization between the complexity control process and the block encoding process. 3. "Pre-data" corresponds to sub-band samples that have been generated by transformation but have not yet been collected across the entire stripe or whose stripes are not yet ready for QP assignment. The height of the pre-data can be understood as the delay between the generation of a new line of sub-band samples and the point at which those samples become active for QP assignment and encoding. A larger delay provides more known statistics for predicting the encoding length of samples existing beyond the active stripe, but this consumes more memory. In many applications, it is desirable to reduce the delay to zero so that there is no pre-data at the time when a new stripe becomes active. 4. "Invisible data" corresponds to sub-band samples that have not yet been generated by transformation.

[0065] As in the case of the complexity control method described in Embodiment 1, for each candidate p of the base Cleanup path (Cup0) bitplane P of the sub-band samples, it is used to generate an estimated value L of the encoding length b Here, the difference is that the estimated value of the encoding length is collected in a record that describes only one stripe of the sub-band. Specifically, the entry L in the record L p (b) k (b) k,p (b)provides a conservative estimate of the number of encoded bytes generated by encoding sub - band samples within stripe k of sub - band b within the HT Cleanup pass for bit - plane p. Generally, these lengths represent multiple code blocks, specifically, all code blocks present within stripe k.

[0066] In some embodiments, the length estimate L k,p (b) can take on fractional or floating - point values rather than being an integer. The method for estimating the encoding length described in Embodiment 2 of the present invention is, of course, adapted to estimate the length contribution from individual line pairs within a sub - band, so it can be useful for calculating and aggregating the estimated values of line - pair lengths with decimal precision. This also forms an estimate of the partial length from any "prior data" as defined above and enables collection within the partial record L adv (b) to be possible.

[0067] In this embodiment, QP values are generated for the active stripe k within sub - band b without waiting for all invisible data to become available. The method for generating these dynamic QP values is substantially the same as that described in Embodiment 1. The main differences are as follows. 1. It is necessary to predict the encoding length associated with sub - band samples that exist beyond the active stripe. 2. It is necessary to track the estimated number of bytes associated with the base Cleanup pass of the code blocks belonging to the dispatched stripe, because they can have different QP values and the base bit - plane P b has been previously committed based on this QP value.

[0068] As shown in FIG. 4, the first problem described here is a prediction length record Λ for representing the estimated lengths associated with all samples beyond the active stripe k k (b)is addressed by a "length predictor" that generates. The prediction process will be further described below.

[0069] The second problem is that when the base bit plane P b is determined using a QP value (equivalently, an integer value F such that QP = F / G), the estimated length B k (b) is extracted from the record L k (b) This is done by the box shown by reference number 10 in the figure, obtaining P b for the code blocks within stripe k using Equation (2) and reporting it as follows.

Number

[0070] These B k (b) values are accumulated for all dispatched stripes of all sub-bands to track the total number of committed bytes B acc Note that B k (b) and B acc are scalar values, while L k (b) and Λ k (b) are vector-valued records representing multiple hypotheses p for the base bit planes that have not yet been determined for the active stripe k.

[0071] The QP estimation procedure is the same as that used in Embodiment 1, except that the target maximum number of bytes L max is decreased by the number of bytes B acc that have already been committed, and Equation (4) becomes as follows.

Number

Number

Number

Number

Number

Number

[0072] Specifically, the extension and bias operations of Equation (3) are as follows here.

Number

Number

Number

[0073] The QP (i.e., F) assignment procedure can be executed whenever a new active stripe becomes available for any sub-band. In this case, the second sum in Equation (8) may include only the sub-band b for which the assignment is being performed, and all other sub-bands may have had their latest active stripes dispatched previously. However, this procedure can also be executed less frequently, waiting until several sub-bands have active stripes ready for QP assignment, and as a result, the second sum in Equation (8) includes multiple terms. Executing the QP assignment procedure less frequently reduces the overall calculations associated with extending, biasing, and accumulating the estimated length records and prediction length records, but this calculation does not become an excessive burden.

[0074] Now, pay attention to the creation of the prediction length record Λ k (b) As shown in Figure 4, the information available for generating predictions consists of the set of active and past estimated length records {L i (b)} 0≦i≦k along with any partial estimated length records L adv (b) already formed from some or all of the prior data.

[0075] Denote N k (b) as the number of sub-band lines represented by these length records, and H (b) as the height of the sub-band, and a simple prediction method is set as follows. [Number]

[0076] In some embodiments, this simple uniform average may be replaced by a weighted average that places more emphasis on more recent estimated length records, which is the number H of subband lines to which the prediction is applied (b) -N k (b) than N k (b) and may be beneficial when large.

[0077] First, if the transform generated only a small number of subband samples, some subbands may not have yet accumulated active stripes, whether dispatched or not. For these subbands b, since there is no latest active stripe, k b = -1. It is important that the first sum in Equation (8) includes predictions from all subbands. In some embodiments, this requirement can be addressed by waiting until all subbands have an active stripe before dispatching the first active stripe from any subband. However, this can consume significant memory resources in deep DWT hierarchies. A preferred approach is to generate the first predicted length record Λ adv (b) as soon as all subbands have accumulated an estimated value L of the partial coding length -1 (b) This can be done after the generation of a single line pair for the subband, and thus the first execution of the QP assignment procedure is delayed until all subbands have received at least one line pair.

[0078] In some embodiments, in order to further reduce latency and memory consumption, an initial prediction for sub-bands within the deep DWT hierarchy can be generated before any data from the transform becomes available. This can be done by scaling the predictions generated by other higher-resolution sub-bands according to the associated sampling density. Fortunately, the low-resolution sub-bands where sub-band samples may not be available when the QP assignment procedure is first executed tend to have a very low sample density and thus have little impact on the vectors used in the QP assignment of Equation (7).

Number

[0079] In other embodiments, a "background" estimated length record L containing the length estimated offline from other image data can be used for some or all of the sub-bands. An initial prediction can then be generated from this background data as follows. bg (b) where W

Number

[0080] In the above description, it may seem that the QP selection is a centralized process that uses synchronized information from all sub-bands to form and distribute decisions for use in determining the base bit-plane value P for block encoding. However, the QP selection process can actually be performed distributively using information that is not necessarily synchronized. In particular, each sub-band or group of sub-bands can be assigned its own local copy of the box indicated by reference numeral 10 in FIG. 4, and these boxes accumulate the committed bytes B b and determine the QP (equivalently, F) value. In a distributed implementation, each such local instance of the committed byte accumulator and QP generation process still requires input from all other sub-bands, but this input can be delayed or partially pre-aggregated for the active stripe for which the QP estimate is being actively generated. In particular, the only external inputs required for the correct operation of the local QP generation process are the cumulative sum of all committed bytes from the external sub-band b, the extended prediction vector that accumulates the latest predictions from the external sub-band Λ k (b) and any length estimate L from external active stripes that have not yet been committed k (b) k (b) .

[0081] Note that the complexity constraint method here does not include the encoding process itself, and thus concludes the description of this Embodiment 3. Unlike the adaptive quantization method employed by many conventional coders, the actual encoding length is not used to determine the QP value. Further, the block encoder is the bit-plane P b ​Since it generates not only the base Cleanup path that depends on QP but also the Z-1 additional encoding path, the QP value itself does not directly determine the quantized subband sample values. These characteristics mean that the block encoding process can be significantly decoupled from the complexity control procedure, enabling an implementation that supports the considerable parallelism provided by independent block encoding. Furthermore, the PCRD-opt algorithm can freely optimize the trade-off of distortion lengths associated with individual code blocks at any point. In some embodiments, the PCRD-opt procedure is only executed when all the encoded data for an image or frame has been generated, which can maximize the opportunity to unevenly distribute bits according to the complexity of the scene. In other embodiments, the PCRD-opt procedure may be executed incrementally to gradually output the final code stream content and reduce memory consumption and latency. However, even in that case, the frequency at which the PCRD-opt procedure is executed can be very different from the frequency at which the QP assignment procedure is executed because the two are separated.

[0082] All of these characteristics and opportunities ultimately stem from the fact that the coding length estimation process is not used for rate control, as previously explained in detail. Embodiment 4: Enhanced Prediction for Video Applications

[0083] In video applications, previously compressed frames can contribute to predicting the coding lengths of invisible data within the current frame. As a starting point, the background length record L bg (b) described above can be derived from the estimated length records within subband b of the previously compressed frame, and this background information can be used not only to form the initial prediction length record Λ -1 (b) but also the normal prediction length record Λ k (b)It can also contribute. Embodiment 4 further provides a method of determining the reliability of the length estimate from the previous frame and incorporating the length estimate from the previous frame into the predicted length within the current frame based on this reliability.

[0084] In this Embodiment 4, a set of summary length records P of J "previous frames" j (b) is maintained for each sub-band b, where 0 ≤ j < J and record j summarizes the coded length estimated for the sub-band in the previous frame for H j (b) lines. As a recommended example, summary records with J = 6 are maintained for each sub-band, and its height H j (b) is roughly divided as follows for the overall height H of the sub-band (b) .

Number

[0085] Using the length estimation method described in Embodiment 2, the estimated value of the coded length is formed from the quad significance statistic values that can become available after each pair of sub-band lines is generated. In this case, the exact summary record height H j (b) must be a multiple of 2. These incremental length estimates are aggregated to form a summary length record C j (b) of the "current frame", which becomes the summary length record P j (b) of the "previous frame" in the next frame. In a memory-efficient embodiment, the C j (b) record can overwrite the P j (b) record as soon as the record is fully generated. For simplicity, the summary record height H j (b) is interpreted as being consistent between frames. Therefore, C j (b) and P j(b) Both represent the estimated coding lengths for the same set of sub - band lines. However, variations of the method that can vary the height can be easily developed.

[0086] N k (b) Recall that N is the number of lines from sub - band b that have been used to form the estimated coding length within the current frame. C jk (b) Let j indicate the next summary - length record C that is currently being assembled. k Let C be the number of summary - length records that have been completed. j (b) Thus, it becomes as follows.

Number

Number

[0087] The set vector of the estimated coding lengths for all N k (b) lines from sub - band b seen so far is as follows,

Number

[0088] A similar vector representing the same number of sub - band lines in the previous frame can be formed as follows.

Number

[0089] Next, for H that does not currently have a length estimate value in the current frame (b) -N k (b) For the subband line, the vector representing the length estimate value in the previous frame can be formed as follows.

Number

[0090] In this embodiment of the present invention, the reliability of the inter-frame length estimate value is compared with the reliability of the intra-frame length estimate value through two quantities.

Number

Number

[0091] Δ temporal (b) <Δ spatial (b) If it is, the prediction vector Λ k (b) is preferably set to P post (b) This is basically assuming that the estimated number of bytes for the missing subband line in the current frame is the same as that for the same subband line in the previous frame, and this is called "temporal prediction". Alternatively, as in Embodiment 3, using Equation (9), C (b) -N k (b) It is preferable to generate Λ pre (b) from, and this is called "spatial prediction". k (b)

[0092] ​ In a preferred embodiment of the present invention, when a sub - band has at least one active stripe within the current frame, extreme pure temporal prediction or extreme pure spatial prediction is avoided by making the following assignments.

Number

[0093] Here,

Number

Number

Number

Number

[0094] Still \(\Delta\) temporal (b) and \(\Delta\) spatial (b) It will be apparent to those skilled in the art that there are many different ways to advantageously perform temporal or spatial prediction using the relationship between them while avoiding both extremes of pure temporal prediction and pure spatial prediction.

[0095] Also, the \(U(P\) pre (b) ) value is the \(P\) that is about to be overwritten by the completed \(C\) j (b) record j (b)It is also clear that it can be incrementally formed by applying Equation (10) to each of the records and accumulating the results. This means that there is no need to provide storage for the summary records of both C j (b) and P j (b) in both sub-bands.

[0096] Furthermore, since the above equations Δ temporal (b) and Δ spatial (b) include the height ratio, but we are only interested in determining which of Δ temporal (b) and Δ spatial (b) is smaller, it will be clear that the costly division operations in these ratios can be avoided by standard multiplication techniques. Embodiment 5: QP Adaptation for Low-Latency Image and Video Encoding

[0097] The foregoing embodiments support high throughput of images and videos by accurate targeting of the encoding length target for each frame of the image or video, i.e., high-quality rate control. Embodiments 3 and 4 of the present invention support a low-memory configuration in which neither the image nor the sub-band samples need to be fully buffered in memory before block encoding. In all cases, the PCRD-opt stage of the entire encoding process can be deferred until all block encodings of the image or video frame (even a group of video frames) are complete, and the encoded bits can be distributed unevenly in space (or even in time) according to the complexity of the scene. In many cases, this is an excellent strategy because the total amount of encoded data that needs to be buffered before the PCRD-opt stage is usually much less than the amount of image or sub-band data it represents.

[0098] In low-latency applications such as sub-frame latency video coding, it is not possible to wait until all the encoded data for an image or frame has been generated before performing the PCRD-opt process and outputting the final code stream. Furthermore, in such applications, it is often necessary to consider a communication channel with a fixed or at least restricted bitrate when determining the end-to-end latency. In the context of JPEG2000, a natural way to address such applications is to collect stripes of code blocks from each subband within an image or video frame into "flush sets" as shown in Figure 5, such that each subband is vertically divided into the same number of flush sets, and each flush set contains contributions from each subband that advance the encoded image representation in a consistent manner. In the case of very low latency, the number of vertical decomposition levels in the wavelet transform is often limited to just 2 or 3, and rectangular code blocks that are much wider than they are tall (e.g., 1024×4) are used. The JPEG2000 precinct dimensions can be chosen to ensure that the height of the code blocks is halved at each level of the DWT hierarchy, and the spatially oriented scan order is chosen such that the encoded information for each flush set can be output to the code stream as soon as it becomes available. The so-called Position, Component, Resolution, Layer (PCRL) scan order should normally be used for low latency, where the encoded data appears in a vertically spatially advancing order (top to bottom), for each spatial position, the image components (usually color planes) appear in order, for each component at each spatial position, successive resolutions appear in order, and successive quality layers (if multiple) for each segment appear in succession. It is also possible to construct flush sets using vertical tiling, but introducing tile boundaries can degrade the properties of the DWT, reduce the coding efficiency, and introduce visual artifacts in the decoded image at low bitrates, so it is a less desirable approach.

[0099] Embodiment 1 is easily adapted to such low-latency encoding environments by restricting the QP assignment process to only subband samples and code blocks belonging to a single flash set. In a constant bitrate or constrained bitrate environment, the number of encoded bytes generated for any given flash set generally needs to satisfy both a lower bound (underflow constraint) and an upper bound (overflow constraint). The upper bound on the compressed size of a flash set becomes the L parameter in Figure 3. When all subband lines of a flash set have been generated by transformation, an estimated encoding length record for the flash set is created using the method described in Embodiment 2, and this is used to assign QP values for that flash set only, based on its L constraint. This QP (equivalently, F) value is used to derive the base bitplane P of all code blocks within the flash set belonging to subband b, and block encoding is performed. Finally, the PCRD-opt algorithm is applied to the generated encoding path to produce a rate-distortion optimal representation of the flash set that conforms to the upper bound. If the generated content cannot meet the lower bound (underflow constraint) of the flash set, stuffing bytes can be inserted into the code block byte stream in a way that does not affect the decoded result, and both the J2K-1 and HT block encoding algorithms support the introduction of stuffing bytes into the encoded content. max When all subband lines of a flash set have been generated by transformation, an estimated encoding length record for the flash set is created using the method described in Embodiment 2, and this is used to max assign QP values for that flash set only, based on its L constraint. This QP (equivalently, F) value is used to derive the base bitplane P of all code blocks within the flash set belonging to subband b, and block encoding is performed. b Finally, the PCRD-opt algorithm is applied to the generated encoding path to produce a rate-distortion optimal representation of the flash set that conforms to the upper bound. max If the generated content cannot meet the lower bound (underflow constraint) of the flash set, stuffing bytes can be inserted into the code block byte stream in a way that does not affect the decoded result, and both the J2K-1 and HT block encoding algorithms support the introduction of stuffing bytes into the encoded content.

[0100] To further reduce latency and / or memory consumption, some embodiments can use the spatial prediction method described in Embodiment 3 or (in the case of video) the combined spatial and temporal prediction method described in Embodiment 4 to enable the block encoding process within a flash set to start before all subband samples of the flash set have been generated by transformation. Experimental Results

[0101] Here, the inventors provide some experimental evidence regarding the effectiveness of the embodiments of the present invention.

[0102] First, consider the compression of a single very large image. The image in question is an aerial photograph with a size of 13333×13333, where each RGB pixel is 24 bits (8 bits / sample) and occupies 533 MB on the disk.

[0103] For the purpose of optimizing the PCRD-opt procedure, the image is compressed to various bitrates using the Mean Squared Error (MSE) and Peak Signal-to-Noise Ratio (PSNR). The code block size is 64×64, and a CDF9 / 7 wavelet transform is used together with a normal irreversible inverse correlation color conversion (in this case, from RGB to YCbCr). The Peak Signal-to-Noise Ratio (PSNR) is selected as the optimization objective so that the effect of the complexity-constrained encoding can be confirmed by simply measuring the PSNR of the decompressed code stream with reference to the original image. In any case, only the HT block coding algorithm is used to compress the image to generate an HTJ2K code stream compliant with JPEG2000 Part 15. All compression and decompression are performed using the methods described so far.

[0104] Compare the performance associated with "Full" HT encoding in which the block encoder starts from the first (coarsest) significant bitplane of each code block and generates a number of passes, with the performance obtained using the method described in Embodiment 3 with minimum latency (no "pre-data" when the stripe becomes active) and various numbers of HT encoding passes Z. The length estimate itself is formed using Embodiment 2. Report the PSNR results for these different approaches in Table 1. Clearly, for at least this type of content, it is sufficient for the HTJ2K encoder to generate a maximum of Z = 6 encoding passes per code block, which is far fewer than the number of passes generated by "Full" HT encoding.

[0105] Table 1: Compression of a large (0.5 GB) RGB aerial image into an HTJ2K code stream using various levels of HT encoder complexity control, corresponding to 3, 4, and 6 encoding passes per code block, followed by unweighted (PSNR-based) PCRD-opt rate control.

Table 1

[0106] Next, the effectiveness of the spatial and temporal composite prediction method described in Embodiment 4 is examined in comparison with the spatial-only (within-frame) prediction method described in Embodiment 3. To do this, an artificial 4K 4:4:4 RGB video sequence consisting of 48 frames with five scene cuts (six segments) that alternate between two different types of content is constructed. Four of the six segments are much easier to compress than the other two, and there is a significant amount of overcast in the upper half of the picture, resulting in a large variation in scene complexity from the top to the bottom of the frame. The content is compressed at a bit rate of 1 bit / pixel (bpp) with constant rate control so that each frame has an essentially the same compression size of 3840×2160×(1 / 8)=1,036,800 bytes. Each frame is encoded to produce an HTJ2K code stream compliant with JPEG2000 Part 15 using only the HT block coding algorithm. The PSNR trace from the mean squared error of each frame across the R, G, and B channels is plotted in FIG. 6.

[0107] The "PCRD-STATS" trace in this figure corresponds to the existing "PCRD-STATS" method for HTJ2K coding with complexity constraints as mentioned in the "Background Art" as described in International Publication No. 2020 / 073098, where the set of encoding paths generated for a given code block is based on the operating point selected by the PCRD-opt procedure for the same code block in the previous frame, and additional encoding paths are introduced to be able to adapt gradually to changes in scene content over time without generating encoding paths with Z>6 for each code block. First, different processing is done for the first frame, and in particular, for the results plotted in FIG. 6, all possible encoding paths are generated for each code block of the first frame.

[0108] The "CPLEX-S" trace in the figure corresponds to the method described in Embodiment 3 and also has the minimum memory, and the "CPLEX-ST" trace corresponds to the method described in Embodiment 4. In both cases, the length estimate itself is formed from the quad significance statistic according to Embodiment 2, and an encoding path with a maximum of Z = 6 is generated for each code block.

[0109] The "HTFULL" trace in the figure corresponds to "Full" HT coding, and many coding paths are generated for each code block, starting from the first (coarsest) bitplane where any sample of the code block is significant. This method is inherently more complex (lower throughput) than other methods because the HT block encoder generates more coding paths. The PCRD-opt procedure is configured to maximize the PSNR (minimize the MSE), and a larger set of options (coding paths) is provided to determine the optimized truncation point of each generated code stream, so "Full" HT coding is also expected to generate the highest PSNR. However, the "Full" HT coding method here uses the gradient threshold prediction function of Kakadu (trademark) software, and the distortion long gradient threshold of the frame is estimated from that used in the previous frame. The estimated gradient is used not to generate all possible coding paths but, if necessary, to terminate the block coding process early. This strategy for complexity control is described in "Software architectures for JPEG2000" by D. Taubman, Proceedings of the IEEE International Conference on DSP, Santorini, Greece (2002), and has been successfully used by JPEG2000 encoders for many years, but as evidenced by trace reference number 20 in the figure, it shows that it is disadvantageous in the transition from frames with high scene complexity to low scene complexity (from low PSNR to high PSNR). (Note: Kakadu software is widely used to implement a series of JPEG2000 standards, and its tools are often used to generate reference results for JPEG2000 in academic and commercial settings. These tools are available at http: / / www.kakadusoftware.com.)

[0110] Note that the "PCRD-STATS" trace (reference number 21 in the figure) shows significant performance loss during scene transitions, especially from frames with low scene complexity to frames with high scene complexity (from high PSNR to low PSNR). This is because it assumes that the complexity of the local scene does not change too rapidly between frames. Therefore, in frames with high complexity following frames with low complexity, many coarse bit planes may be skipped.

[0111] The pure intra-frame complexity control method from Embodiment 3 (the "CPLEX-S" curve of reference number 22) is more temporally robust than both "HTFULL" and "PCRD-STATS", but it suffers some quality degradation in the easily compressible parts of videos where the upper part of each frame has very low scene complexity. Embodiment 3 of the present invention uses the estimated value of the coding length from the upper part of each subband already generated by spatial transformation to generate the predicted length of the remaining (invisible) part of the subband, so this is not surprising.

[0112] Overall, the combined spatial and temporal prediction method of Embodiment 4 (the "CPLEX-ST" curve of reference number 23) is superior to all other methods in both temporal robustness and overall compressed image quality.

[0113] The encoding method of the above-described embodiments can be implemented by a suitable computing device programmed with suitable software. Embodiments of the method can be implemented in GPU configurations as well as in low-latency and high-latency hardware configurations.

[0114] When software is used to implement the embodiments, the software can be provided on a computer-readable medium such as a disk, or as a data signal on a network such as the Internet, or in any other way.

[0115] The above embodiments relate to use within the JPEG2000 (HTJ2K) and J2K formats. Embodiments of the present invention are not limited thereto. Some embodiments may be used with other image processing formats. Embodiments can find applications in other image processing contexts.

[0116] Those skilled in the art will understand that numerous variations and / or modifications can be made to the present invention as shown in the specific embodiments without departing from the spirit or scope of the invention as broadly described. Therefore, the present embodiment should be considered illustrative in all respects and not restrictive.

[0117] In the following claims and the foregoing description of the invention, unless the context requires otherwise, the word "comprise", or variations of that word such as "comprises" or "comprising", is used in an inclusive sense, that is, to specify the presence of the stated features but not to preclude the presence or addition of further features in various embodiments of the invention.

[0118] It should be understood that if any prior art publication is referred to herein, such reference does not admit that the publication forms part of the common general knowledge in the art in Australia or any other country.

Claims

Claim 1 A method for constrained encoding of a code stream, including a code stream of JPEG2000 and high-throughput JPEG2000, subject to a target full-length constraint, comprising: a. collecting information regarding subband samples generated by spatial transformation; b. generating an estimated value of the encoding length from the collected information for a plurality of potential bit-plane truncation points; c. determining a quantization parameter (QP value) from these generated estimated values of the encoding length, the QP value being associated with one of the potential bit-plane truncation points, and after mapping the QP value to the base bit-plane index of each associated code block of each subband, when subsequent encoding step e is executed, determining the QP value such that the overall encoding length estimated when truncating at these bit-plane indexes does not exceed the target full-length constraint; d. mapping the QP value to the base bit-plane index of each such code block; e. encoding each associated code block with an accuracy associated with the corresponding base bit-plane index, and encoding one or more additional encoding passes from each code block; f. subjecting all the encoding passes thus generated to a post-compression rate-distortion optimization process to determine a final set of encoding passes output from each code block as a compression result. Claim 2 The method according to claim 1, wherein the block encoding algorithm of JPEG2000 Part 1 is used. Claim 3 The method according to claim 1, wherein the high-throughput block encoding algorithm of JPEG2000 Part 15 is used, and the base bit-plane for a code block corresponds to the first high-throughput (HT) Cleanup pass generated for that code block. Claim 4 Before determining the global QP value, an estimated value of the coding length is formed for all samples in each subband, and then the base bitplane index for all code blocks is determined. Subsequently, based on these indexes, the coding of the associated coding path can be advanced. The method according to any one of claims 1 to 3.

5. The method according to any one of claims 1 to 3, wherein an estimated value of the coding length is incrementally formed as subband samples become available from the spatial transform.

6. The QP value is incrementally updated according to a coding length budget that starts in a state equal to the target total length constraint. As code blocks become incrementally available for coding, the code blocks are identified as "dispatched code blocks". Subsequently, a. collecting the length estimate values based on the samples available in the subband into an estimated length record; b. generating a predicted length record for predicting the coding length estimated for samples that are not yet available in the subband; c. incrementally determining the QP value based on the estimated length record and the predicted length record. The QP value is associated with one of the potential bitplane truncation points. After mapping the QP value to the base bitplane index for each non-dispatched code block in each subband, when the subsequent coding step e is executed, the estimated total coding length of the non-dispatched code block is expected not to exceed the remaining length budget when truncated at these bitplane indexes. The step of incrementally determining the QP value; d. mapping the QP value to the base bitplane index of the non-dispatched code block that is identified as an "active code block" when the subband samples are available; e. inputting the active code block into a set of dispatched code blocks and subtracting the estimated coding length corresponding to the truncation at each base bitplane index from the coding length budget. f. encoding the dispatched code blocks to a precision associated with each base bitplane index and encoding one or more additional encoding passes from each such code block; g. making all such generated encoding passes available for a post-compression rate distortion optimization process, the method of claim 5. **Claim 7** The method according to any one of claims 1 to 6, wherein an estimated value of the encoding length for the collection of subband samples is formed using a simplified model of a block encoding process lacking one or more features of the actual encoding process, and the estimated length for a given bitplane truncation point is almost surely greater than the actual encoding length at that same point. **Claim 8** The method of claim 7, wherein the estimated value of the encoding length for the collection of subband samples is formed from significance statistics, and the samples are significant at the bitplane boundary if the quantized magnitude in that bitplane is non-zero. **Claim 9** The method of claim 8, wherein the estimated value of the encoding length for the collection of subband samples is formed from quad significance numbers, the quads consist of four samples, and the number for a given bitplane identifies the number of quads from said collection that will contain at least one significant sample within that bitplane. **Claim 10** The method of claim 6, wherein the length prediction is obtained by extrapolating an estimated length determined from available subband samples within the same subband. **Claim 11** The method of claim 6 or claim 10, wherein the prediction incorporates background information regarding typical estimated lengths of previously collected similar image content. **Claim 12** The method of claim 6, wherein the complexity-constrained encoding procedure is applied to video, and the length prediction for a given subband within the current video frame is determined using subband samples observed in the previous frame. **Claim 13** The method according to claim 12, wherein a length prediction for samples not yet available in a given subband of the current video frame is formed by combining a length estimate value formed for corresponding subband samples in the previous frame, which is identified here as a temporal prediction, with an extrapolated length estimate value formed from available subband samples in the current frame, which is identified here as a spatial prediction.

14. The method according to claim 13, wherein the temporal and spatial predictions are combined based on a reliability measure that compares the consistency of length estimate values of available subband samples across frames with the consistency of length estimate values of subband samples that are respectively available and unavailable in the current frame and in the previous frame.

15. The code blocks of an image or video frame are partitioned into flash sets such that each flash set has its own coding length constraint, and all of the QP generation, base bitplane mapping, block coding, and post-compression rate distortion optimization processes operate on a per-flash-set basis based on the individual flash set length constraints, the method according to any one of claims 1 to 14.

16. An apparatus for constrained coding of a code stream, the apparatus comprising an encoder configured to implement the method according to any one of claims 1 to 15.

17. A computer program comprising instructions for controlling a computer to implement the method according to any one of claims 1 to 15.

18. A non-volatile computer-readable medium providing the computer program according to claim 17.

Citation Information

Patent Citations

  • A method and apparatus for image compression

    WO2017201574A1

  • A further improved method and apparatus for image compression

    WO2019191811A1