Implementation of image coding hardware
Patent Information
- Application Number
- BR112025020723
- Authority / Receiving Office
- BR · BR
- Patent Type
- Applications
- Publication Date
- 2026-08-25
Smart Images

Figure 00000000_0000_ABST
Description
1 / 35 “IMAGE CODING HARDWARE IMPLEMENTATION” TECHNICAL FIELD
[0001] This disclosure relates to a hardware implementation for converting input data elements in a first format to output data elements in a second format. In particular, but not exclusively, the invention relates to a hardware implementation for converting input data elements representing image data in a first format to output data elements representing image data in a second format. BACKGROUND
[0002] Image encoding processes can be implemented in software or hardware, and while software implementation is relatively simple, hardware implementation offers significant advantages in terms of speed and performance.
[0003] A hardware implementation of image encoding can provide faster encoding and decoding of image data compared to a software implementation. This is because the hardware implementation can be specifically designed to perform specific image encoding tasks and optimized to provide faster processing times. For many use cases, a basic or obvious hardware implementation is sufficient, but for high-demand applications, such as encoding high-resolution or high-frame-rate images, more advanced hardware implementations are required.
[0004] However, implementing image encoding in hardware also brings its own set of challenges. High-demand applications often translate into high costs in terms of resources and processing time, as the hardware needs to be able to handle large amounts of data. Furthermore, image encoding often involves using transformation operations to transform image data from one format to a different format. The type of transformation required can vary. Petition 870250107895, dated 11 / 25 / 2025, page 9 / 56 2 / 35 between frames and even within a frame, which makes creating a hardware implementation capable of performing these transformations challenging. To address these challenges, there is a need for an optimized hardware implementation that can efficiently perform image encoding at reduced costs and with the ability to handle multiple transformations.
[0005] Another challenge in hardware implementation is that high-resolution, high-frame-rate applications require significant resources, both in terms of area and processing capacity. This can result in high device costs, and in many cases, the cost becomes prohibitive for low-cost devices. To overcome this challenge, it is necessary to design an optimized hardware implementation that can efficiently perform image encoding using minimal resources while simultaneously meeting performance requirements. SUMMARY
[0006] According to a first aspect of the invention, a hardware block is provided for converting input data elements representing image data in a first format to output data elements representing image data in a second format. The hardware block comprises a plurality of modules organized into a first group of modules and a second group of modules. Each module in the first group of modules is configured to: receive a subset of the input data elements in the first format; and perform a plurality of operations to convert the subset of input data elements in the first format into intermediate elements.Each intermediate element is derived from the subset of input data elements in the first format using one of several operations; and where each module of the second group of modules is configured to: receive a subset of the intermediate elements, where each intermediate element of the subset received from the intermediate elements is from a different module of the first group of modules; and perform the various operations to convert the subset of intermediate elements into a subset of the... Petition 870250107895, dated 11 / 25 / 2025, p. 10 / 56 3 / 35 output data elements in the second format.
[0007] Preferably, when multiple modules are configured to perform the same operations.
[0008] Preferably, based on an indicator received in the hardware block, intermediate elements derived from at least one module of the first group of modules are emitted by the hardware block without being sent to the second group of modules.
[0009] Preferably, where each intermediate element is derived from the subset of input data elements in the first format, using an operation distinct from plurality of operations.
[0010] Preferably, where each subset received from the intermediate elements is derived from the subset of the input data elements in the first format using the same distinct operation of plurality of operations.
[0011] Preferably, where the input data elements in the first format are residual elements and the output data elements in the second format are a set of transformed elements indicative of an extension of the spatial correlation in the residual elements.
[0012] Preferably, wherein the output data elements in the second format are residual elements and the input data elements in the first format are a set of transformed elements indicative of an extension of the spatial correlation in the residual elements.
[0013] Preferably, where the set of transformed elements indicates one or more of the average, horizontal, vertical, and diagonal relationships between neighboring residual elements.
[0014] Preferably, where the set of transformed elements is based on a direct Hadamard decomposition transformation.
[0015] Preferably, where the residual elements are based on a difference between a first rendering of an image associated with the image data at a quality level in a hierarchy in Petition 870250107895, dated 11 / 25 / 2025, page 11 / 56 4 / 35 layers with various quality levels and a second rendering of the image at the same quality level.
[0016] Preferably, where each of the various operations performs a multidimensional direct Hadamard decomposition transformation.
[0017] Preferably, wherein the subset of input data elements in the first format and the subset of output data elements in the second format each comprise four data elements.
[0018] Preferably, wherein the first group of modules and the second group of modules comprise four modules each.
[0019] Preferably, where the first group of modules and the second group of modules are the same.
[0020] Preferably, wherein the intermediate element derived in at least one module of the first group of modules is emitted by the hardware block.
[0021] Preferably, wherein the intermediate element derived from at least one module of the first group of modules is emitted by the hardware block in response to the hardware block receiving an indicator.
[0022] Preferably, the indicator can be derived from information contained in a bit stream received in the hardware.
[0023] Preferably, where the bitstream comprises input data elements and metadata.
[0024] Preferably, metadata comprises the information used to derive the indicator.
[0025] According to a second aspect of the invention, a method is provided for converting input data elements representing image data in a first format to output data elements representing image data in a second format. The method comprises the use of a plurality of modules organized into a first group of modules and Petition 870250107895, dated 11 / 25 / 2025, page 12 / 56 5 / 35 a second group of modules. The method further comprises, in each module of the first group of modules: receiving a subset of the input data elements in the first format; and performing a plurality of operations to convert the subset of input data elements in the first format into intermediate elements. Each intermediate element is derived from the subset of input data elements in the first format using one of several operations. The method further comprises, in each module of the second group of modules: receiving a subset of the intermediate elements, wherein each intermediate element of the subset received from the intermediate elements is from a different module of the first group of modules; and performing the plurality of operations to convert the subset of intermediate elements into a subset of the output data elements in the second format.
[0026] Preferably, where each intermediate element is derived from the subset of input data elements in the first format, using an operation distinct from plurality of operations.
[0027] Preferably, where each subset received from the intermediate elements is derived from the subset of the input data elements in the first format using the same distinct operation of plurality of operations.
[0028] Preferably, where the input data elements in the first format are residual elements and the output data elements in the second format are a set of transformed elements indicative of an extension of the spatial correlation in the residual elements.
[0029] Preferably, wherein the output data elements in the second format are residual elements and the input data elements in the first format are a set of transformed elements indicative of an extension of the spatial correlation in the residual elements.
[0030] Preferably, where the set of transformed elements indicates one or more of the average, horizontal, vertical, and diagonal relationships between neighboring residual elements. Petition 870250107895, dated 11 / 25 / 2025, page 13 / 56 6 / 35
[0031] Preferably, where the set of transformed elements is based on a direct Hadamard decomposition transformation.
[0032] Preferably, wherein the residual elements are based on a difference between a first rendering of an image associated with the image data at a quality level in a layered hierarchy with multiple quality levels and a second rendering of the image at the same quality level.
[0033] Preferably, where each of the various operations performs a multidimensional direct Hadamard decomposition transformation.
[0034] Preferably, where the subset of input data elements in the first format and the subset of output data elements in the second format each comprise four data elements.
[0035] Preferably, where the first group of modules and the second group of modules are the same.
[0036] Preferably, where the intermediate element derived in at least one module of the first group of modules is emitted by the hardware block.
[0037] Preferably, wherein the intermediate element derived from at least one module of the first group of modules is emitted by the hardware block in response to the hardware block receiving an indicator.
[0038] Preferably, the indicator can be derived from information contained in a bit stream received in the hardware.
[0039] Preferably, the bitstream comprises input data elements and metadata.
[0040] Preferably, metadata comprises the information used to derive the indicator. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The invention will now be described, only by way of Petition 870250107895, dated 11 / 25 / 2025, page 14 / 56 7 / 35 example, with reference to the attached drawings, in which:
[0042] Figure 1 shows an example of how data coded for a Quality Level - LOQ-1 - is generated in a coding device;
[0043] Figure 2 shows an example of how LOQ-0 is generated in an encoding device;
[0044] Figure 3 schematically shows an example of how the decoding process is performed;
[0045] Figure 4 illustrates two directional decomposition transformations (e.g., Hadamard) in a conversion process;
[0046] Figure 5 is a block diagram showing an exemplary hardware arrangement for implementing the DD transformation or the iDD transformation of Figure 4;
[0047] Figure 6 illustrates the possible hardware blocks that can be used to perform the operation of converting residual data to the AHVD AA coefficient;
[0048] Figure 7 illustrates the possible hardware blocks that can be used to perform the operation of converting residual data to the AHVD AH coefficient;
[0049] Figure 8 illustrates possible hardware blocks that can be used to perform the operation of converting AHVD coefficients into a residual data element;
[0050] Figure 9 illustrates the possible hardware blocks that can be used to perform the operation of converting the AHVD_4x4 coefficients into another residual data element;
[0051] Figure 10 illustrates a modular design of a dd_common module;
[0052] Figure 11 illustrates the dd_common modules arranged to perform a DDs transformation; and
[0053] Figure 12 illustrates the arrangement of the modules Petition 870250107895, dated 11 / 25 / 2025, page 15 / 56 8 / 35 dd_common from Figure 11 in an example where an iDDs transformation is performed. DETAILED DESCRIPTION
[0054] The hierarchical frame encoding methods described herein include generating residues for a full frame, then a decimated frame, and so on. Different levels in the hierarchy may relate to different resolutions, here termed Quality Levels - LOQs - and residual data may be generated for different levels. In some examples, the residual data from video compression for a full-size video frame may be termed “LOQ-0” (e.g., 1920x1080 for a High Definition - HD video frame), while the residual data from the decimated frame may be termed “LOQ-x”. In these cases, “x” indicates the number of hierarchical decimations. In some examples described in this document, the variable “x” has a maximum value of one, and therefore there are exactly two hierarchical levels for which compression residues will be generated (e.g., x = 0 and x = 1).
[0055] Figure 1 shows an example of how data coded for a Quality Level - LOQ-1 - is generated in a coding device.
[0056] The algorithm and general methods are described using an AVC / H.264 encoding / decoding algorithm as an example of a baseline algorithm. However, other encoding / decoding algorithms can be used as baseline algorithms without any impact on the operation of the general algorithm. Figure 1 shows the 100 process of generating entropy-encoded residues for the LOQ-1 hierarchical level.
[0057] The first step 101 is to decimate an uncompressed input video by a factor of two. However, other factors can be used to decimate and / or scale the uncompressed input video. This may involve reducing the sampling rate of an input frame 102 (labeled “Input Frame” in Figure 1) with height H and width W to generate a frame Petition 870250107895, dated 11 / 25 / 2025, page 16 / 56 9 / 35 decimal 103 (labeled as “Half 2D size” in Figure 1) with height H / 2 and width W / 2. The sampling rate reduction process involves reducing each axis by a factor of two and is effectively performed through the use of 2x2 grid blocks. Sampling rate reduction can be done in several ways, including, but not limited to, averaging and Lanczos resampling.
[0058] The decimated frame 103 then goes through a base coding algorithm (in this example, an AVC / H.264 coding algorithm) where an entropy-encoded reference frame 105 (called “2D Half-Size Base” in Figure 1) with height H / 2 and width W / 2 is generated by an entity 104 called “H.264 Encoding” in Figure 1 and stored as H.264 entropy-encoded data. However, other scaling factors can be used, depending on the scaling mode. The entity 104 may comprise a coding component of a base encoder-decoder, for example, a base codec or a base encoding / decoding algorithm. A base-encoded data stream can be output as an entropy-encoded reference frame 105, where the base-encoded data stream is at a lower resolution than an input data stream that provides the input frame 102.
[0059] In the present example, an encoder simulates a decoding of the output of entity 104. A decoded version of the encoded reference frame 105 is then generated by an entity 106 called “H.264 Decoding” in Figure 1. Entity 106 may comprise a decoding component of a base codec. The decoded version of the encoded reference frame 105 may represent a version of the decimated frame 103 that would be produced by a decoder after receiving the entropy-encoded reference frame 105.
[0060] In the example in Figure 1, a difference is calculated between the decoded reference frame issued by entity 106 and the decimated frame 103. This difference is referred to in this document as “LOQ-1 residuals”. A Petition 870250107895, dated 11 / 25 / 2025, page 17 / 56 The 10 / 35 difference forms an input to a transformation block 107.
[0061] The transformation (in this example, a Hadamard-based transformation) used by transformation block 107 converts the difference into four components. Transformation block 107 can perform a directed (or directional) decomposition to produce a set of coefficients or components related to different aspects of a set of residuals. In Figure 1, transformation block 107 generates the coefficients A (mean), H (horizontal), V (vertical), and D (diagonal). Transformation block 107, in this case, exploits the directional correlation between the LOQ-1 residuals, which was found to be surprisingly effective, in addition to, or as an alternative to, performing a transformation operation to a higher quality level – an LOQ-0 level. The LOQ-0 transformation is described in more detail below.In particular, it was identified that, in addition to exploiting directional correlation in LOQ-0, directional correlation may also be present and surprisingly effectively exploited in LOQ-1 to provide more efficient coding than exploiting directional correlation only in LOQ-0 or not exploiting directional correlation in LOQ-0 at all.
[0062] The coefficients (A, H, V, and D) generated by the transformation block 107 are then quantized by a quantization block 108. Quantization can be performed using variables called “step widths” (also called “step sizes”) to produce quantized transformed residues 109. In one example, each quantized transformed residue 109 has a height H / 4 and a width W / 4. For example, if a 4x4 block of an input frame is taken as a reference, each quantized transformed residue 109 could have a pixel height and width. However, other scaling factors can be used, depending on the scaling mode. Quantization involves reducing the decomposition components (A, H, V, and D) by a predetermined factor (step width). The reduction can be performed by division, for example, dividing the coefficient values by a step width, for example, representing a bin width for quantization. Quantization can generate a Petition 870250107895, dated 11 / 25 / 2025, page 18 / 56 11 / 35 set of coefficient values with a range of values smaller than the range of values that enter the quantization block 108 (for example, transformed values in a range of 0 to 21 can be reduced using a step width of 7 to a range of values between 0 and 3). In a hardware implementation, an inverse of a set of step width values can be precalculated and used to perform the reduction via multiplication, which can be faster than division (for example, multiplying by the inverse of the step width).
[0063] The quantized residues 109 are then entropy encoded to remove any redundant information. Entropy encoding may involve, for example, passing the data through a run-length encoder (RLE) 110 followed by a Huffman encoder 111.
[0064] The encoded and quantized components (Ae, He, Ve, and De) are then placed into a serial stream with definition packets inserted at the beginning of the stream. The definition packets may also be called header information. Definition packets may be inserted frame by frame. This final stage may be performed using a file serialization routine 112. The definition packet data may include information such as the Huffman encoder specification 111, the type of sample rate boost to be employed, whether or not the A and D coefficients will be discarded, and other information that allows the decoder to decode the streams. The residual output data 113 is therefore entropy-encoded and serialized.
[0065] Both the reference data 105 (the mid-size baseline entropy-encoded frame) and the entropy-encoded LOQ-1 residual data 113 are generated for decoding by the decoder during a reconstruction process. In one case, the reference data 105 and the entropy-encoded LOQ-1 residual data 113 can be stored and / or buffered. The reference data 105 and the entropy-encoded LOQ-1 residual data 113 can be communicated to a decoder for decoding. Petition 870250107895, dated 11 / 25 / 2025, page 19 / 56 12 / 35
[0066] In the example in Figure 1, several additional operations are performed to produce a set of residues at another quality level (e.g., higher) - LOQ-0. In Figure 1, several decoder operations for the LOQ-1 stream are simulated in the encoder.
[0067] First, the quantized output 109 is branched and reverse quantization 114 (or “dequantization”) is performed. This generates a representation of the coefficient values generated by the transformation block 107. However, the representation generated by the dequantization block 109 will be different from that generated by the transformation block 107, as there will be errors introduced due to the quantization process. For example, several values in a range of 7 to 14 can be replaced by a single quantized value of 1 if the step width is 7. During dequantization, this single value of 1 can be dequantized by multiplying itself by the step width to generate a value of 7. Therefore, any value in the range of 8 to 14 will introduce an error in the output of the dequantization block 109.Since the highest level of LOQ-0 quality is generated using dequantized values (e.g., including a simulation of decoder operation), LOQ-0 residuals can also encode a correction for quantization / dequantization error and any errors due to downsampling / upsampling operations that may remove certain frequency components from the image, depending on the filter characteristics associated with the downsampling / upsampling operations.
[0068] Secondly, an inverse transformation block 115 is applied to the unpacked coefficient values generated by the dequantization block 114. The inverse transformation block 115 applies a transformation that is the inverse of the transformation performed by the transformation block 107. In this example, the transformation block 115 performs an inverse Hadamard transformation, although other transformations could be used. The inverse transformation block 115 converts the dequantized coefficient values (e.g., values of A, H, V, and D in a coding block or unit) back into corresponding residual values (e.g., Petition 870250107895, dated 11 / 25 / 2025, page 20 / 56 13 / 35 representing a reconstructed version of the input of the transform block 107). The output of the inverse transform block 115 is a set of reconstructed LOQ-1 residues (e.g., representing an output from a decoding process of the LOQ-1 decoder). The reconstructed LOQ-1 residues are added to the decoded reference data (e.g., the output of the decoding entity 106) to generate a reconstructed video frame 116 (labeled “Half-size 2D Recon (For LOQ-0)” in Figure 1) with height H / 2 and width W / 2 (other scaling factors may be used depending on the scaling mode). The reconstructed video frame 116 closely resembles the originally decimated input frame 103, as it is reconstructed from the output of the decoding entity 106, but with the addition of the reconstructed LoQ-1 residues. The reconstructed video frame 116 is a temporary output for a LOQ-0 mechanism.This process mimics the decoding process, and that is why the originally decimated frame 103 is not used. The addition of the reconstructed LOQ-1 residuals to the decoded basestream, i.e., the output of the decoding entity 106, allows the LOQ-0 residuals to also correct errors introduced into the LOQ-1 stream by quantization (and, in certain cases, by transformation), for example, as well as errors related to the reduction and increase of the sampling rate.
[0069] Figure 2 shows an example of how LOQ-0 is generated 200 in an encoding device.
[0070] To obtain the LOQ-0 residues, the 216-size LOQ-1 reconstructed frame (labeled “2D Half-Size Recon (of LOQ-1)” in Figure 2) is derived as described above with reference to Figure 1. For example, the 216-size LOQ-1 reconstructed frame comprises the 116-frame reconstructed video.
[0071] The next step is to perform a scaling up of the reconstructed 216 frame to full size, WxH. In this example, the scaling up is a factor of two. At this point, several algorithms can be used to improve the scaling up process, such as the nearest, bilinear, acute, or cubic algorithms. The frame Petition 870250107895, dated 11 / 25 / 2025, page 21 / 56 The reconstructed 14 / 35 full-size frame 217 is referred to as the “Predicted Frame” in Figure 2, as it represents a prediction of a frame with full width and height, as decoded by a decoder. The reconstructed full-size frame 217, with height H and width W, is subtracted from the original uncompressed video input 202, creating a set of residues, referred to here as “LOQ-0 residues”. LOQ-0 residues are created at a higher quality level (e.g., resolution) than LOQ-1 residues.
[0072] Similar to the LOQ-1 process described above, the LOQ-0 residuals are transformed by a transformation block 218. This may involve the use of a directed decomposition, such as a Hadamard transformation, to produce coefficients or components A, H, V, and D. The output of transformation block 218 is then quantized by means of quantization block 219. This may be performed based on defined step widths, as described for the first quality level (LOQ-1). The output of quantization block 219 is a set of quantized coefficients, and in Figure 2, these are entropy-encoded 220, 221 and serialized into a file 222. Again, entropy encoding may involve the application of run-length encoding 220 and Huffman encoding 221. The output of entropy encoding is a set of entropy-encoded output residuals 223.They form an LOQ-0 stream, which can be output by the encoder, as well as the LOQ-1 stream (i.e., 113) and the base stream (i.e., 105). The streams can be stored and / or buffered before further decoding by a decoder.
[0073] As can be seen in Figure 2, a “predicted average” component 224 (described in more detail below and denoted Aenc below) can be derived using data from the reconstructed video frame (LOQ-1) 116 before the sample rate upscaling process. This can be used in place of the A (average) component in the transformation block 218 to further improve the efficiency of the encoding algorithm.
[0074] Figure 3 schematically shows an example of how the 300 decoding process is executed. This process of Petition 870250107895, dated 11 / 25 / 2025, page 22 / 56 15 / 35 decoding 300 can be performed by a decoder.
[0075] The 300 decoding process begins with three input data streams. The decoder input therefore consists of entropy-encoded data 305, LOQ-1 entropy-encoded residual data 313, and LOQ-0 entropy-encoded residual data 323 (represented in Figure 3 as serialized encoded data in a file). The entropy-encoded data 305 includes the reduced-size encoded base, for example, data 105, as shown in Figure 1. The entropy-encoded data 305 is, for example, half the size, with dimensions W / 2 and H / 2 relative to the total frame with dimensions W and H.
[0076] The entropy-encoded 305 data is decoded by a base-306 decoder using the decoding algorithm corresponding to the algorithm used to encode that data (in this example, an AVC / H.264 decoding algorithm). This may correspond to decoding entity 106 in Figure 1. At the end of this step, a decoded 325 video frame is produced, with reduced size (e.g., half the size) (indicated in this example as an AVC / H.264 video). This can be viewed as a standard resolution video stream.
[0077] In parallel, the entropy-encoded LOQ-1 313 residual data is decoded. As explained above, the LOQ1 residuals are encoded into four components (A, V, H, and D) which, as shown in Figure 3, have a dimension of one-quarter the dimension of the total frame, i.e., W / 4 and H / 4. This is because, as also described below and in prior patent applications US 13 / 893,669 and PCT / EP2013 / 059847, the contents of which are incorporated herein by reference, the four components contain all the information associated with a given direction within the untransformed residuals (i.e., the components are defined relative to a block of untransformed residuals). As described above, the four components can be generated by applying a 2x2 transformation kernel to the residuals whose dimension, for LOQ-1, would be W / 2 and H / 2, in other words, the same dimension. Petition 870250107895, dated 11 / 25 / 2025, page 23 / 56 16 / 35 of the reduced-size, entropy-encoded data 305. In the decoding process 300, as shown in Figure 4, the four components are entropy-decoded in the entropy decoding block 326 and then dequantized in the dequantization block 314 before an inverse transformation is applied via the inverse transformation block 315 to generate a representation of the original LOQ-1 residuals (e.g., the input of transformation block 107 in Figure 1). The inverse transformation may comprise a Hadamard inverse transformation, for example, as applied to a 2x2 block of residual data. The dequantization block 314 is the inverse of the quantization block 108 described above with reference to Figure 1.At this stage, the quantized values (i.e., the output of entropy decoding block 326) are multiplied by the step width factor (i.e., step size) to generate reconstructed transformed residues (i.e., components or coefficients). It can be observed that blocks 114 and 115 in Figure 1 mirror blocks 314 and 315 in Figure 3.
[0078] The decoded LOQ-1 residues, for example, as output from the inverse transform block 315, are then added to the decoded video frame, for example, the output from the base decoding block 306, to produce a reconstructed video frame 316 at a reduced size (in this example, half the size), identified in Figure 3 as “2D Half Size Recon”. This reconstructed video frame 316 is then subjected to a sample rate boost to achieve full resolution (e.g., the 0th quality level from the 1st quality level) using a sample rate boost filter, such as bilinear, bicubic, sharp, etc. In this example, the reconstructed video frame 316 is subjected to a sample rate boost from half width (W / 2) and half height (H / 2) to full width (W) and full height (H).
[0079] The reconstructed video frame with increased sampling rate 317 will be a predicted frame in LOQ-0 (normal size, WxH) to which the decoded residues in LOQ-0 are added.
[0080] In Figure 3, the residual data encoded in Petition 870250107895, dated 11 / 25 / 2025, page 24 / 56 17 / 35 LOQ-0 323 are decoded using an entropy decoding block 327, a dequantization block 328, and an inverse transformation block 329. As described above, the LOQ-0 323 residual data are encoded using four components (i.e., they are transformed into components A, V, H, and D) which, as shown in Figure 3, have a dimension of half the dimension of the total frame, i.e., W / 2 and H / 2. This is because, as described in this document and in prior patent applications US 13 / 893,669 and PCT / EP2013 / 059847, the contents of which are incorporated herein by reference, the four components contain all the information relating to the residuals and are generated by applying a 2x2 transformation kernel to the residuals whose dimension, for LOQ-0, would be W and H, in other words, the same dimension as the total frame.The four components are entropy decoded by the entropy decoding block 327, then dequantized by the dequantization block 328, and finally transformed 329 back into the original LOQ-0 residues by the inverse transformation block 329, transformation (for example, in this example, a 2x2 Hadamard inverse transformation).
[0081] The decoded LOQ-0 residues are then added to the predicted frame 317 to produce a complete reconstructed video frame 330. Frame 330 is an output frame, with height H and width W. Therefore, the decoding process 300 in Figure 3 is capable of generating two user data elements: a base-decoded video stream 325 at the first quality level (e.g., a half-resolution stream at LOQ-1) and a full-resolution or higher-resolution video stream 330 at a higher quality level (e.g., a full-resolution stream at LOQ-0).
[0082] The above description was made with reference to specific sizes and baseline algorithms. However, the above methods apply to other sizes and / or baseline algorithms. The above description is given only as an example of the more general concepts described in this document.
[0083] Figure 4 illustrates two transformations of Petition 870250107895, dated 11 / 25 / 2025, page 25 / 56 18 / 35 Directional decomposition (e.g., Hadamard) in a conversion process.
[0084] The goal of the conversion process is to convert the residuals into directional decomposed values (direct transformation) and convert the directional decomposed values back into the original residuals (inverse transformation). As mentioned in Figures 1 to 3, the residuals are the values derived from subtracting the reconstructed video frame from the ideal input frame (or with reduced sampling rate).
[0085] First, Figure 4 illustrates a 2x2 residual data block 405 and the corresponding 2x2 AHVD coefficient block 410, as mentioned above. The 2x2 AHVD coefficient block 410 can be derived from the 2x2 residual data block 405 using a direct 2x2 Hadamard transformation DD_2x2 (DD transformation). The 2x2 residual data block 405 is derived from the 2x2 AHVD coefficient block 410 using an inverse or reverse 2x2 Hadamard transformation iDD_2x2 (iDD transformation). The DD transformation in Figure 4 can be used to perform the transformation on blocks 107, 218 in Figures 1 and 2. The iDD transformation in Figure 4 can be used to perform the inverse transformation on blocks 115, 314, 329 in Figures 1 and 3.
[0086] Figure 4 also illustrates a larger 4x4 residual data block, 415, and the corresponding 4x4 AHVD coefficient block 420. The 4x4 AHVD coefficient block 420 is derivable from the 4x4 residual data block 415 using a direct 4x4 Hadamard transformation DD_4x4 (DDs transformation). The 4x4 residual data block 415 is derived from the 4x4 AHVD coefficient block 420 using an inverse or reversed 4x4 Hadamard transformation iDD_4x4 (iDDs transformation). 2x2 Transformations
[0087] To calculate a DD transformation for the 2x2 405 residual data block in Figure 4, the following equations are used: A = / ?oo + R01+ ^o + H = Roo - fyi + *io — *11 Petition 870250107895, dated 11 / 25 / 2025, page 26 / 56 19 / 35 F = Λοο + *01 — *ιο — *ιι D = Ro— *01 — *10 + *11
[0088] For simplicity, the equations above do not include an averaging factor that would be used to avoid incorrect sizing.
[0089] To calculate an iDD transformation for the 2x2 coefficient block AHVD 410 in Figure 4, the following equations are used: οο = + ++ ο1 = - +1ο = + -- = - -+
[0090] For simplicity, the equations above do not include an averaging factor that would be used to avoid incorrect sizing.
[0091] Figure 5 is a block diagram showing an exemplary hardware arrangement 500 for implementing the DD transformation or the iDD transformation of Figure 4 (the DD and iDD transformations in 4x4 are described in relation to Figures 6 to 9). Figure 5 shows four inputs that go into the hardware arrangement to produce four outputs. The four inputs are input 1, input 2, input 3, and input 4, which can be residual data in the form of 00, *01, *10, and Rn respectively for a DD transformation, or they can be AHVD coefficients AB, C, and D respectively for an iDD transformation. The four outputs are output 1, output 2, output 3, and output 4, which can be the AHVD coefficients A, B, C, and D respectively for a DD transformation, or they can be residual data in the form of 00, 01, 10, and 11 respectively for an iDD transformation.
[0092] The 500 hardware arrangement is now described in more detail.
[0093] The addition block 505a adds input 1 and input 2, and the resulting value is stored in register 510a. The subtraction block 505b subtracts input 2 from input 1, and the resulting value is stored in register 510b. Petition 870250107895, dated 11 / 25 / 2025, page 27 / 56 20 / 35 The addition block 505c adds input 3 and input 4, and the resulting value is stored in register 510c. The subtraction block 505d subtracts input 4 from input 3, and the resulting value is stored in register 510d.
[0094] The sum block 515a adds the value stored in register 510a and the value stored in register 510c and stores the resulting value in register 525a after passing through the shift operator 520a. The sum block 515b adds the value stored in register 510b and the value stored in register 510d and stores the resulting value in register 525b after passing through the shift operator 520b. The subtraction block 515c subtracts the value stored in register 510c from the value stored in register 510a and stores the resulting value in register 525c after passing through the shift operator 520c. The subtraction block 515d subtracts the value stored in register 510d from the value stored in register 510b and stores the resulting value in register 525d after passing through the shift operator 520d.The 520a-d shift operators are optionally used to obtain the correct weights / scale of the resulting output components; for example, a division by 4 can be used to provide an average instead of the absolute sum.
[0095] The value stored in register 525a is then forwarded to register 530a and register 535a for storage and output as output 1. The value stored in register 525b is then forwarded to register 530b and register 535b for storage and output as output 2. The value stored in register 525c is then forwarded to register 530c and register 535c for storage and output as output 3. The value stored in register 525d is then forwarded to register 530d and register 535d for storage and output as output 4. These registers can be used to achieve pipelining efficiency.
[0096] One or all of the registers used in Figure 5 may not be used, so that the data from the computation blocks (i.e., blocks of addition, subtraction, and / or shift operators) can be passed from the input to the computation block and directly to the output without being stored in multiple registers. Petition 870250107895, dated 11 / 25 / 2025, page 28 / 56 21 / 35 4x4 Conversions
[0097] To calculate a DDs transformation for the 4x4 residual data block 415 in Figure 4, the following equations are used: AA = 0OO+ «01 + «02 + «3 + ^0 + «n+ 112+ «13 + «20 + «21 + «22 + «23 + «30 + «31 + «32 + «33 AH = «00 + «01-«02-«03 + «10 + «11-«12-«13+«20+«21-«22-«23+«30+«31-«32-«33 4V = «00+«01+«02+«03+«10+«11+«12+«13-«20-«21-«22-«23-«30-«31-«32-«33 = 00 + 01 - 02 - 03 + 10 + 11 - 12 - 13 - 20 - 21 + 22 + 23 - 30-«31+«32+«33 HA = "00 - "01 + "02 - "03 + "10 - "11 + "12 - "13 + "20 - "21 + "22 - "23 + "30-"31+"32-"33 HH = "00 - "01 - "02 + "03 + "10 - "11 - "12 + "13 + "20 - "21 - "22 + "23 + "30-"31-"32+"33 HV = "00 - "01 + "02 - "03 + "10 - "11 + "12 - "13 - "20 + "21 - "22 + "23 - "30 + "31 - "32 + "33 HD = "00 - "01 - "02 + "03 + "10 - "11 - "12 + "13 - "20 + "21 + "22 - "23 - "30+"31+"32-"33 VA = «00+«01+«02+«03-«10-«11-«12-«13+«20+««21+«22+«23-«30-«31-«32-«33 VH = "00+"01-"02-"03-"10-"11+"12+"13+"20+"21-"22-"23-"30-"31+"32+"33 VV = «00+«01+«02+«03-«10-«11-«12-«13-«20-«21-«22-«23+««30+««31+«32+«33 VD = "00 + "01 - "02 - "03 - "10 - "11 + "12 + "13 - "20 - "21 + "22 + "23 + "30+"31-"32-"33 DA = «00-«01+«02-«03-«10+««11-«12+««13+««20-«21+««22-«23-««30+««31-«32+«33 Petition 870250107895, of 25 / 11 / 2025, p. 29 / 56 22 / 35 DH - T?oo - R01- R02+ / ^3 - T^o + / / ,, + R12- ^3 + R20- R21- R22+ R23- R30+ *31+*32-*33 DV - Roo- T?01 + *02 - *03 - *10 + *11 - *12 + *13 - *20 + *21 - *22 + *23 + *30-*31+*32-*33 DD - «00 - *01 - *02 + *03 - *10 + *11 + *12 - *13 - *20 + *21 + *22 - *23 + *30-*31-*32+*33
[0098] For simplicity, the equations above do not include an averaging factor that would be used to avoid incorrect sizing.
[0099] To calculate an iDDs transformation for the 4x4 AHVD 420 coefficient block in Figure 4, the following equations are used: - + + + + + + + + + + + ++ ++ - + - - + + - - + + - - ++ -02 - - + - + - + - + - + - ++03 - - - + + - - + + - - + +-+ - + + + + + + + - - - - -- -- - + - - + + - - - - + + -- ++ - - + - + - + - - + - + -+ -+ - - - + + - - + - + + - -+ +- - + + + - - - - + + + + -- -Petition 870250107895, of 25 / 11 / 2025, p. 30 / 56 23 / 35 R21= AA + AH - AV - AD - HA- Η H + HV + HD + VA + VH - VV - VD - DA - DH + DV + DD R22= AA —AH + AV -AD -HA + Η Η -Η V + HD + VA - VH + VV -VD-D A + DH -+ R23= AA -AH -AV + A - — HA + HH + HV -HD + V— -V—-VV + V --DA + DH +R30= AA + AH + AV + AD - HA - HH - HV - HD - VA - VH - VV - VD + DA + DH + DV + DD = + - - + - + - + + + -- = - + - - + - + - + - + +- +- = - - + - + + - - + + - +- -+
[0100] For simplicity, the above equations do not include an averaging factor that would be used to prevent incorrect sizing.
[0101] Figure 6 illustrates the possible hardware blocks that can be used to perform the operation of converting residual data to the AHVD AA coefficient. Residual data is received in receiving block 605. From receiving block 605, the residual data is summed in computing block 610 to calculate the AHVD AA coefficient. Although, in this example, the residual data is initially received in receiving block 605 before being passed to computing block 610, in other examples, receiving block 605 may not be used and the residual data may be received directly in computing block 610.
[0102] Figure 7 illustrates the possible hardware blocks that can be used to perform the operation of converting residual data to the AHVD AH coefficient. Residual data is received in receiving block 705. From receiving block 705, residual data 10, 11, 12, Petition 870250107895, dated 11 / 25 / 2025, page 31 / 56 24 / 35 R 13, ^30, 31±, ^32e33s are subtracted from the sum of the residual data Z?00, T?01, T?02, T?01, R 20 > #21, ^22e / ?21 to calculate the AHVD coefficient AA. Although, in this example, the residual data is initially received in receiving block 705 before being passed to computing block 710, in other examples, receiving block 705 may not be used and the residual data may be received directly in computing block 710.
[0103] Figure 8 illustrates the possible hardware blocks that can be used to perform the operation of converting AHVD_4x4 coefficients into residual data00, in other words, to execute part of an iDDs. The hardware shown in Figure 8 is the hardware in Figure 6. The AHVD_4x4 coefficients are received in the receiving block 805. From the receiving block 805, the AHVD_4x4 coefficients are summed in the computing block 810 to calculate the residual data00. Although, in this example, the AHVD_4x4 coefficients are initially received in the receiving block 805 before being passed to the computing block 810, in other examples, the receiving block 805 may not be used and the AHVD_4x4 coefficients may be received directly in the computing block 810.
[0104] Figure 9 illustrates the possible hardware blocks that can be used to perform the operation of converting the AHVD_4x4 coefficients into residual data01. The hardware shown in Figure 9 is the same hardware as in Figure 7. The AHVD_4x4 coefficients are received in the receiving block 905. In the receiving block 905, the AHVD_4x4 coefficients , , , , , , and are subtracted from the sum of the AHVD_4x4 coefficients AA, AH, HA, HH, VA, VH, DA and DH to calculate the residual data T?01. Although, in this example, the AHVD_4x4 coefficients are initially received in receive block 905 before being passed to compute block 910, in other examples, receive block 905 may not be used and the AHVD_4x4 coefficients may be received directly in compute block 910.
[0105] To convert the AHVD_4x4 coefficients into a Petition 870250107895, dated 11 / 25 / 2025, page 32 / 56 25 / 35 block 4x4 of residual data or vice versa, 16 different blocks (similar to those shown in Figures 6 to 9) will be needed to execute the 16 equations shown above to produce residual data / ?00a / ?33 or coefficients AHVD_4x4 AA a DD. Using 16 different blocks to perform 16 different mathematical operations translates into high costs in terms of resources (e.g., physical area for hardware) and processing time. Furthermore, each block needs to perform 15 additions or subtractions. This can make it difficult to complete the operation in each block in a single clock cycle or other desired time constraint. Therefore, it is necessary to optimize the above approach to perform the transformations.
[0106] Another problem also arises based on the type of transformation used (2x2 DD or 4x4 DDs), which can change between frames (and even within a frame); therefore, it is a challenge to provide an encoder / decoder with the ability to use the same hardware to perform both DD and DD transformations. It is less efficient if the encoder / decoder needs separate hardware to perform 2x2 DD / iDD and 4x4 iDD / DD transformations.
[0107] Therefore, the disclosure contained herein aims to provide a modular piece of hardware that can be used in all implementations of DD, iDD, DDs and iDDs and that can reduce the number of calculations in a single hardware block to process the transformation more quickly.
[0108] The modular hardware component was derived as follows:
[0109] The transformation equations for iDDs are as follows: R 00 = AA + AH + AV + AD + HA + HH + HV + HD + VA + VH + VV + VD + DA + DH + DV + DD R01=AA + AH - AV - A + +H+ + HH - HV - HD + VA + V H-VV-V D + DA + DH -DV-DD Petition 870250107895, of 25 / 11 / 2025, p. 33 / 56 26 / 35 R02= AA — AH + AV - A + + HA — HH + Η V -HD + VA - VH + VV - V + + DA - DH + DV - DD R03= AA —AH—AV + AD +HA - HH -Η V + HD + VA-VH-V + + V+ +DA - DH -D + + DD R10= AA + AH + AV + AD + HA + HH + HV + H-- (VA + VH + VV + VD) - (DA + DH + DV + DD) R 11 = AA + AH-AV-AD + HA + HH - HV - HD - (VA + VH-VV- VD) - (DA + DH -DV - DD) = - + - + - + - - ( - + -) - ( -DH + DV - DD) = - - + + - - + - ( - - +) - ( -DH -DV + DD) = + + + - ( + + + ) + + + +- ( + DH + DV + DD) = + - - - ( + - - ) + + - -- ( + DH -DV -DD) = - + - - ( - + - ) + - + -- ( -DH + DV -DD) = - - + - ( - - + ) + - - +- ( -DH -DV + DD) = + + + - ( + + + ) - ( + + +) + +++ = + - - - ( + - - ) - ( + - -) + +-- = - + - - ( - + - ) - ( - + -) + -DH + D V-DD = - - + - ( - - + ) - ( - - +) + --+
[0110] If you analyze the signs of the equations above, you can Petition 870250107895, dated 11 / 25 / 2025, p. 34 / 56 27 / 35 distinguish a consistent pattern of algebraic operations that can be simplified even further: J = A + H + V + DK=A+HVD L=A-H+VD M = A-Η-+ + D
[0111] Thus, the system of equations above, which represents all the residuals, can finally be expressed as follows: R 00 = AJ + HJ + VJ + DJ Rio = AJ + HJ — VJ — DJ R 20 =AJ — HJ + V-DJ R 3o =AJ-HJ-VJ + D] R01 = AK + HK + VK + DK R 11 = AK + HK — VK — DK R21= A - — HK + V --DK = - -+ = + ++ = + -22 = - +- = - -+ = + ++13= + -- = - +- = - -+
[0112] By analyzing the equations presented so far, it can be observed that there is a repeated pattern of operations to calculate the iDDs transformation, which mimics the iDD transformation, and which can be exploited to apply modularity and reuse of blocks in the project.
[0113] For example, compare Petition 870250107895, dated 11 / 25 / 2025, p. 35 / 56 28 / 35 J = + + H + V + DK= Α +H — V — DL=Α-Η+--D =--+
[0114] with R00 = A + H + V + D R01= - — Η + VD = + - - = - - +
[0115] Thus, the iDD and iDDs transformations can be reduced to common terms for the modularized architecture. Furthermore, the terms involved in the iDDs can be rearranged and then expressed as expressions that correspond to the output of the DD / iDD.
[0116] The repeated pattern of operations for computing iDDs, which mimics the iDD transformation, can be exploited to apply modularity and block reuse in the design. The modular implementation advantageously allows the reuse of the same structure to calculate the 2x2 and 4x4 transformations.
[0117] Although this has been shown above for the relationship between iDDs and iDD, the same applies to the relationship between DDs and DD.
[0118] Figure 10 illustrates a modular design of a dd_common 1000 module. Figure 10 shows four inputs that go into the hardware arrangement to produce four outputs. The four inputs are input 1, input 2, input 3, and input 4, which can be residual data for a direct transformation, or they can be AHVD coefficients for an inverse or reverse transformation. The four outputs are output 1, output 2, output 3, and output 4, which can be AHVD coefficients for a direct transformation or they can be residual data for an inverse or reverse transformation.
[0119] Input register 1005 receives input 1, input 2, input 3, and input 4. Addition block 1010a adds input 1 and input 2, and the resulting value is stored in register 1015a. The subtraction block Petition 870250107895, dated 11 / 25 / 2025, page 36 / 56 29 / 35 1010b subtracts input 2 from input 1, and the resulting value is stored in register 1015b. The addition block 1010c adds input 3 and input 4, and the resulting value is stored in register 1015c. The subtraction block 1010d subtracts input 4 from input 3, and the resulting value is stored in register 1015d.
[0120] The sum block 1020a adds the value stored in register 1015a and the value stored in register 1015c and forwards the resulting value to output register 1025 to provide output 1. The sum block 1020b adds the value stored in register 1015b and the value stored in register 1015d and forwards the resulting value to output register 1025 to provide output 2. The subtraction block 1020c subtracts the value stored in register 1015c from the value stored in register 1015a and forwards the resulting value to output register 1025 to provide output 3. The subtraction block 1020d subtracts the value stored in register 1015d from the value stored in register 1015b and forwards the resulting value to output register 1025 to provide output 4.
[0121] The registers used in Figure 10 are optional and may not be used, so that the data from the computation blocks (i.e., addition and subtraction blocks) are passed from input to output without being stored in the registers.
[0122] Figure 11 illustrates the dd_common modules arranged to perform a DDs transformation. Figure 11 shows 8 dd_common modules (reference 1000 in Figure 10) to perform the DDs transformation, which is half the number of modules / blocks that would be needed if the arrangement of Figures 6 to 9 were used. Each individual dd_common module performs a 2x2 DD transformation, but the combination of all modules allows the execution of a 4x4 DDs transformation.
[0123] The arrangement of the dd_common module in Figure 11 is organized in two stages to ensure efficient calculation of 4x4 transformations. The arrangement in Figure 11 comprises a first stage (stage 0) and a second stage (stage 1). Stage 0 comprises 4 dd_common modules and stage 1 comprises 4 dd_common modules. Each of the modules in stage 0 is connected Petition 870250107895, dated 11 / 25 / 2025, page 37 / 56 30 / 35 to each of the modules in stage 1, as can be seen in Figure 11. Each dd_common module in stage 0 receives 4 inputs and produces 4 intermediate values. In this example, the 4 inputs for each dd_common module in stage 0 are residual values to be transformed. Each dd_common block in stage 1 receives the 4 intermediate values emitted by the dd_common blocks in stage 0 and outputs the AHVD coefficients. The arrangement also allows the computation of 2x2 transformations, as will be discussed later.
[0124] In the example in Figure 11, the inputs to stage 0 are the 4x4 residuals (i.e., T?00 for / ?33 data). The outputs 1 to 4 of stage 0, the intermediate values, are labeled x00, x01, x10, and x11, respectively (where x denotes the first letter of the two-letter form of the second-order transformation elements (AHVD_4x4) which are ultimately the output of stage 1. For example, in Figure 11, a first module dd_common in stage 1 is arranged to be x=A, so that its inputs are intermediate values A00, A01, A10, and A11). Stage 1 generates the transformation coefficients AHVD_4x4 (i.e., AA to DD). Therefore, the final AHVD_4x4 transformation coefficients are obtained at the output of stage 1 (second stage), taking into account the outputs of the first stage.
[0125] A discussion of a 4x4 transformation is given below in more detail with reference to Figures 10 and 11.
[0126] The inputs / ? 00, Λ01, / ? 10 and / ? 11 are received in a first dd_common module of the dd_common module array in stage 0 as inputs 1-4 respectively, to produce intermediate outputs A00, H00, V00 and D00 as outputs 1-4 respectively. The intermediate values A00, H00, V00 and D00 act as input 1 for the respective dd_common modules of stage 1.
[0127] The inputs / ? 02, T? 03, / ? 12 and / ? 13 are received in a second dd_common module of the dd_common module array in stage 0 as inputs 1-4 respectively, to produce intermediate outputs A01, H01, V01 and D01 as outputs 1-4 respectively. The intermediate values A01, H01, V01 and Petition 870250107895, dated 11 / 25 / 2025, p. 38 / 56 31 / 35 D01 acts as input 2 for the respective dd_common modules of stage 1.
[0128] The inputs f? 20, / ? 21, / ? 30 and f? 31 are received in a third dd_common module of the dd_common module array in stage 0 as inputs 1-4 respectively, to produce intermediate outputs A10, H10, V10 and D10 as outputs 1-4 respectively. The intermediate values A10, H10, V10 and D10 act as input 3 for the respective dd_common modules of stage 1.
[0129] The inputs / ? 22, / ? 23, / ? 32 and f? 33 are received in a fourth dd_common module of the dd_common module array in stage 0 as inputs 1-4 respectively, to produce intermediate values A11, H11, V11 and D11 as outputs 1-4 respectively. The intermediate outputs A11, H11, V11 and D11 act as input 4 for the respective dd_common modules of stage 1.
[0130] In stage 1, a fifth dd_common module (the first dd_common module in stage 1 mentioned above) receives intermediate values A00, A01, A10 and A11 as inputs 1-4, respectively, to produce outputs AA, AH, AV and AD as outputs 1-4, respectively.
[0131] In stage 1, a sixth dd_common module receives the intermediate values H00, H01, H10 and H11 as inputs 1-4, respectively, to produce the outputs HA, HH, HV and HD as outputs 1-4, respectively.
[0132] In stage 1, a seventh dd_common module receives the intermediate values V00, V01, V10 and V11 as inputs 1-4, to produce the outputs VA, VH, VV and VD as outputs 1-4, respectively.
[0133] In stage 1, an eighth dd_common module receives intermediate values D00, D01, D10 and D11 as inputs 1-4, respectively, to produce outputs DA, DH, DV and DD as outputs 1-4, respectively.
[0134] In this example, each intermediate value received in a stage 1 dd_common module is received from a different stage 0 dd_common module. However, in other examples, multiple intermediate values received in a stage 1 dd_common module are received from the same stage 0 dd_common module.
[0135] In the case of a 2x2 transformation being Petition 870250107895, dated 11 / 25 / 2025, page 39 / 56 32 / 35 executed, the outputs of each dd_common module in stage 0 would be the AHVD_2x2 coefficients obtained as a result of a DD transformation operation, for example, A = Roo+ R01+ R10+ R1±, H = R00-R01+ Λ10-Λ11, = = Roo+ R01R10- and = = Roo- / ?01- / ?10+ Rt11. When executing a 2x2 transformation, the outputs of one or each of the dd_common modules in stage 0 are directly emitted (not shown in Figure 11) for further processing or storage as AHVD_2x2 coefficients. Not all modules in stage 0 need to be used; For example, only the dd_common module associated with Λ00, T?01, T?10, and Λ11 can be used to produce the corresponding AHVD_2x2 coefficients for the output. Using all dd_common modules from stage 0 to perform separate 2x2 transformations in parallel is advantageous, at least in terms of efficient hardware usage.In one example, a stage 0 dd-common module can perform a 2x2 transformation on its input and send the results directly for further processing or storage (optionally without sending the outputs to stage 1 dd-common modules) based on an indicator. Optionally, the indicator can be received in the hardware block or the relevant stage 0 dd-common module with the input data or separately. Optionally, the indicator can be obtained from a received bitstream (in some embodiments, it can be a bitstream).
[0136] The arrangement of multiple dd_common modules as illustrated in Figure 11 enables a system that can be used for DD and DD transformations and minimizes hardware resources by reducing DD and DD transformations to common terms for a modularized architecture.
[0137] Figure 12 illustrates the arrangement of the dd_common modules from Figure 11 in an example where an iDDs transformation is performed. To perform an inverse transformation, the same arrangement of the dd_common module used in Figure 11 can be used.
[0138] In the example in Figure 12, the inputs for stage 0 are “second-order transformation elements” or AHVD_4x4 coefficients. Petition 870250107895, dated 11 / 25 / 2025, pp. 40 / 56 33 / 35 (i.e., data AA to DD). The outputs of stage 0 / input of stage 1 (i.e., the intermediate values) are xJ, xL, xK, and xM (where x denotes the first letter of the two-letter form of the second-order transformation elements; for example, in Figure 12, the upper left corner or the first dd_common has x=A, therefore it produces AJ, AL, AK, and AM). The result of stage 1 is the residual data (i.e., 00^33).
[0139] In the case of a 2x2 inverse transformation being performed, the inputs to stage 0 are “first-order transformation elements” or AHVD_2x2 coefficients (i.e., given A, H, V, and D) and outputs 1-4 of each dd_common module are 4 residual values, respectively, for example R00=A+H+V+ D, / ? 01= A -H + - - D, R10=A + H -V -D and Λ11 = ---- V + .
[0140] As can be seen in the description of Figures 11 and 12, a 4x4 transformation requires two banks of four dd_common modules, while a 2x2 transformation is obtained with a single dd_common module (as discussed in Figures 11 and 12). A dd_common module / block is defined as a module configured to receive 4 inputs and produce 4 outputs, where, if the 4 inputs are first-order transformation elements, such as A, H, V, D, the 4 outputs will be a 2x2 block of residues, where, if the 4 inputs are part of a group of second-order transformation elements, such as AHVD_4x4 elements AA, AH, etc., arranged as in Figure 12, the output at stage 0 will be a set of intermediate values and the subsequent stage, stage 1, will be used to obtain and produce the 4x4 residual values. Clearly, if the inputs to stage 0 are 4x4 residual values, the outputs in stage 1 will be AHVD_4x4 transformation elements.
[0141] The new dd_common module and arrangement disclosed here increase throughput and reduce hardware components.
[0142] Instead of a single stage of 16 computing blocks, the invention uses a first stage with 4 modules connected to a Petition 870250107895, dated 11 / 25 / 2025, pp. 41 / 56 34 / 35 second stage with 4 modules. In this example, each module is modular (i.e., it processes the data in the same way). An inventive aspect of the invention is how to route / connect each output of each of the modules in the first stage (“intermediate output”) to the “correct” module in the second stage to obtain the desired final outputs.
[0143] The intermediate results (i.e., the results of the first stage) correspond to the final result of the DD / iDD transformations, which means that a single architecture can be used for DD / iDD transformations and for DDs / iDDs transformations.
[0144] The throughput is increased because (in the more complex example, iDDs or DDs) each final output takes only the time of 4 additions (although there are 8 additions in each module, they are performed as 4 executed in parallel and then another 4 in parallel, therefore, in terms of processing time, it is the same as 2 additions within the dd_common block. Then there are two module stages. Thus, the total “time” is the same as 4 additions) instead of (a single set of) 15 additions in the examples in Figures 6 to 9, which would require the total “time” of 15 additions. The more elements are added, the more time is required, therefore, a reduction from 15 to 4 reduces the time to generate the result. A shorter time per output increases productivity because more outputs can be produced in a given time.Instead of using a single set of 15 additions to process 16 inputs, as shown in the examples in Figures 6 to 9, a 16-input tree structure can be used, which may comprise four stages of additions in this case (i.e., A+B, C+D, etc., resulting in 8 outputs for stage 1, followed by 4 outputs in stage 2, and so on). Using a tree structure would result in an increased throughput, but with the disadvantage of increasing the number of logic resources used. However, the disclosure contained herein allows for an increase in throughput and a drastic reduction in the logic resources required.
[0145] Now analyzing the calculation of the entire transformation Petition 870250107895, dated 11 / 25 / 2025, pp. 42 / 56 35 / 35 iDDs or DDs require 8 dd_common modules, which in turn require 64 additions (8 additions in each dd_common module, and there are 8 dd_common modules). The additions are executed 16 in parallel, followed by 16 in parallel to complete the stage 0 additions, so in terms of processing time, it's the same as 2 additions. Taking into account the stage 1 modules leads to a total processing time equal to 4 additions. Conversely, to compute the entire iDDs or DDs transformation, 16 blocks would be needed, each with 15 additions (totaling 240 additions). However, the additions of the 16 blocks can be executed in parallel, resulting in a total "time" of 15 additions to compute the entire iDDs or DDs transformation.
[0146] Furthermore, the hardware components are reduced because, for the more complex computation (iDDs or DDs transformations), only 8 adder blocks (dd_common modules) are needed, instead of the 16 adder blocks needed in the examples in Figures 6 to 9. If the stage 0 adder modules are reused, only 4 adder modules are needed in total to compute the iDDs or DDs transformations. By reusing stage 0 modules as stage 1 modules, the output of the module set when they are in stage 0 is routed back to the said module set to undergo the subsequent stage 1 operation.
[0147] The above embodiments should be understood as illustrative examples. Other embodiments are provided. It should be understood that any feature described in relation to any embodiment may be used alone or in combination with other described features, and may also be used in combination with one or more features of any other embodiment, or any combination of any other embodiment. Furthermore, equivalents and modifications not described above may also be employed without departing from the scope of the invention, which is defined in the appended claims. Petition 870250107895, dated 11 / 25 / 2025, pp. 43 / 56
Claims
1 / 5 CLAIMS 1. Hardware block for converting input data elements representing image data in a first format to output data elements representing image data in a second format; wherein, the hardware block is characterized in that it comprises: a plurality of modules organized into a first group of modules and a second group of modules, wherein the first group of modules and the second group of modules comprise four modules each, wherein each module of the first group of modules is configured to: receive a subset of the input data elements representing image data in the first format;and perform a plurality of operations to convert the subset of input data elements representing image data in the first format into intermediate elements, wherein each intermediate element is derived from the subset of input data elements representing image data in the first format using one of the pluralities of operations; and wherein each module of the second group of modules is configured to: receive a subset of the intermediate elements, wherein each intermediate element of the subset received from the intermediate elements is from a different module of the first group of modules; and perform the plurality of operations to convert the subset of intermediate elements into a subset of the output data elements representing image data in the second format.
2. Hardware block, according to claim 1, characterized in that each intermediate element is derived from the subset of input data elements representing the image data in the first format using a distinct operation from the plurality of operations. Petition 870250087410, dated 09 / 26 / 2025, pp. 114 / 120 2 / 5 3. Hardware block, according to claim 2, characterized in that each received subset of the intermediate elements is derived from the subset of input data elements representing image data in the first format using the same distinct plurality operation.
4. A hardware block, according to any of the preceding claims, characterized in that the input data elements representing the image data in the first format are residual elements, and the output data elements representing the image data in the second format are a set of transformed elements indicative of an extension of the spatial correlation in the residual elements.
5. Hardware block, according to any one of claims 1 to 3, characterized in that the output data elements representing the image data in the second format are residual elements and the input data elements representing the image data in the first format are a set of transformed elements indicative of an extension of the spatial correlation in the residual elements.
6. Hardware block, according to any one of claims 4 or 5, characterized in that the set of transformed elements indicates one or more average, horizontal, vertical and diagonal relationships between neighboring residual elements.
7. Hardware block, according to claim 6, characterized in that the set of transformed elements is based on a direct Hadamard decomposition transformation.
8. Hardware block, according to any one of claims 4 to 7, characterized in that the residual elements are based on a difference between a first rendering of an image associated with the image data at a quality level in a layered hierarchy with multiple quality levels and a second rendering of the image at the same quality level. Petition 870250087410, dated 09 / 26 / 2025, pp. 115 / 120 3 / 5 9. A hardware block, according to any of the preceding claims, characterized in that each of the various operations performs a multidimensional Hadamard direct decomposition transformation.
10. Hardware block, according to any of the preceding claims, characterized in that the subset of input data elements representing image data in the first format and the subset of output data elements representing image data in the second format each comprise four data elements.
11. Hardware block, according to any of the preceding claims, characterized in that the first group of modules and the second group of modules are the same.
12. Hardware block, according to any of the preceding claims, characterized in that the intermediate element derived in at least one module from the first group of modules is issued by the hardware block.
13. Method for converting input data elements representing image data in a first format to output data elements representing image data in a second format; wherein the method is characterized by the fact that it comprises the use of a plurality of modules organized into a first group of modules and a second group of modules, wherein the first group of modules and the second group of modules each comprise four modules, wherein the method further comprises in each module of the first group of modules: receiving a subset of the input data elements representing image data in the first format; and performing a plurality of operations to convert the subset of input data representing image data elements in the first format into intermediate elements, wherein each intermediate element is derived from the subset of input data elements that Petition 870250087410, dated 09 / 26 / 2025, p.116 / 120 4 / 5 represent image data in the first format using one of the plurality of operations; and wherein the method further comprises each module of the second group of modules: receiving a subset of the intermediate elements, wherein each intermediate element of the subset received from the intermediate elements is from a different module of the first group of modules; and executing the plurality of operations to convert the subset of intermediate elements into a subset of the output data elements that represent image data in the second format.
14. Method, according to claim 13, characterized in that each intermediate element is derived from the subset of input data elements representing the image data in the first format using an operation distinct from the plurality of operations.
15. Method, according to claim 14, characterized in that each subset received from the intermediate elements is derived from the subset of input data elements that represent image data in the first format using the same distinct plurality operation.
16. A method, according to any one of claims 13 to 15, characterized in that the input data elements representing the image data in the first format are residual elements and the output data elements representing the image data in the second format are a set of transformed elements indicative of an extension of the spatial correlation in the residual elements.
17. Method, according to any one of claims 13 to 15, characterized in that the output data elements representing the image data in the second format are residual elements and the input data elements representing the image data in the first format are a set of transformed elements indicative of an extension of the spatial correlation in the residual elements. Petition 870250087410, dated 09 / 26 / 2025, pp. 117 / 120 5 / 5 18. A method, according to any one of claims 16 or 17, characterized in that the set of transformed elements indicates one or more average, horizontal, vertical, and diagonal relationships between neighboring residual elements.
19. Method, according to claim 18, characterized in that the set of transformed elements is based on a direct Hadamard decomposition transformation.
20. A method, according to any one of claims 16 to 19, characterized in that the residual elements are based on a difference between a first rendering of an image associated with the image data at a quality level in a layered hierarchy with multiple quality levels and a second rendering of the image at the same quality level.
21. A method, according to any of the preceding claims, characterized in that each of the various operations performs a multidimensional direct Hadamard decomposition transformation.
22. A method, according to any of the preceding claims, characterized in that the subset of input data elements representing the image data in the first format and the subset of output data elements representing the image data in the second format each comprise four data elements.
23. A method, according to any of the preceding claims, characterized in that the first group of modules and the second group of modules are the same.
24. A method, according to any of the preceding claims, characterized in that the intermediate element derived in at least one module of the first group of modules is issued by the hardware block. Petition 870250087410, dated 09 / 26 / 2025, pp. 118 / 120