JPEG-AI-based entropy coding hardware accelerator and entropy coding method

By designing an entropy encoding hardware accelerator based on JPEG-AI, the problem of inefficient image encoding caused by the inability to be performed by the CPU in the JPEG AI standard is solved, and efficient image compression and real-time requirements are achieved, and image compression of any input resolution and code rate is supported.

CN120343289APending Publication Date: 2025-07-18NANJING UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510575247.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the existing JPEG AI standard, entropy encoding can only be performed through the CPU, resulting in inefficient image encoding and unable to meet real-time requirements.

Method used

A JPEG-AI-based entropy encoding hardware accelerator is designed, including a masking unit, a data loading unit, an on-chip cache unit, an encoding unit and a data write back unit. By reorganizing the storage of the conversion table and the status table, the calculation and access storage time is reduced, and image compression of any input resolution and code rate is supported.

Benefits of technology

It improves the execution efficiency and image encoding efficiency of entropy encoding, can support image compression of various resolutions and compression code rates, reduces calculation and storage time, and meets real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343289A_ABST
    Figure CN120343289A_ABST
Patent Text Reader

Abstract

The invention provides an entropy coding hardware accelerator and entropy coding method based on JPEG-AI.The accelerator comprises a mask unit, a data loading unit, an on-chip cache unit, a coding unit and a data write-back unit, and the coding unit comprises a residual coding module and a super-tensor coding module; the mask unit performs mask operation on the residual tensor and the scale tensor to output a mask residual tensor, a mask scale tensor and effective element information, and the on-chip cache unit stores a recombination conversion table and a recombination state table output by the data loading unit and input and output data of the coding unit. The residual encoding module encodes the mask residual tensor based on the mask scale tensor, the recombination conversion table and the recombination state table and outputs a residual code stream, and the supertensor encoding module encodes the supertensor based on the recombination conversion table so as to output a supertensor code stream. And if the code rate is not changed, the table corresponding to the current code rate stored on the chip can be directly used, so that re-access is avoided, and the image coding efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image compression technology, and particularly to an entropy coding hardware accelerator and an entropy coding method based on JPEG-AI. Background Art

[0002] JPEG AI is an image coding standard based on machine learning. It analyzes and encodes images using a machine learning model to achieve efficient compression. JPEG AI coding first calculates the latent representation of the input image through a neural network, and then compresses the latent representation into a bitstream through entropy coding.

[0003] Entropy coding is a lossless data compression method. By re-encoding the data, the average code length after encoding is made as close as possible to the information entropy of the information source, thereby achieving efficient compression of the data. Since the entropy coding algorithm specified by the JPEG AI standard is different from the methods of various previous entropy coding hardware accelerators, and the CPU can quickly adjust the coding algorithm and parameters according to different requirements, it can only be executed by the CPU.

[0004] Although existing neural network hardware accelerators can effectively accelerate the neural network calculations in the JPEG AI standard, since the entropy coding in the standard can only be executed by the CPU, there is still a large room for accelerating the overall image coding efficiency. Summary of the Invention

[0005] This application provides an entropy coding hardware accelerator and an entropy coding method based on JPEG-AI to solve the problem of low image coding efficiency.

[0006] In a first aspect, this application provides an entropy coding hardware accelerator based on JPEG-AI, including:

[0007] A mask unit for performing a mask operation on the input data to output a mask residual tensor, a mask scale tensor, and valid element information, where the input data includes a residual tensor and a scale tensor;

[0008] A data loading unit for loading a recombined conversion table, a recombined status table, a residual tensor, a scale tensor, and a supertensor from an off-chip memory into an on-chip cache unit, where the recombined conversion table is the recombined conversion table and the recombined status table is the recombined status table;

[0009] An on-chip cache unit for storing the recombined conversion table and the recombined status table;

[0010] An encoding unit including a residual encoding module and a supertensor encoding module. The residual encoding module encodes the mask residual tensor based on the mask scale tensor, the recombined conversion table, the recombined status table, and the valid element information to output a residual bitstream;

[0011] The super-tensor encoding module encodes the super-tensor based on the recombination conversion table to output a super-tensor bitstream;

[0012] A data write-back unit is configured to read the masked residual tensor, masked scale tensor, residual bitstream, and super-tensor bitstream stored in the on-chip cache unit and write them back to the off-chip memory.

[0013] In some feasible embodiments, a data buffer is provided in the masking unit, and the data buffer includes a first data storage unit and a second data storage unit, and the first data storage unit and the second data storage unit are alternately used;

[0014] When performing the masking operation, the masking unit is specifically configured as:

[0015] If an element of the scale tensor is greater than a preset value, mark the element as a true element, and obtain a corresponding element in the residual tensor based on the element to obtain a pair of retained elements, and write the pair of retained elements into the data buffer, where the pair of retained elements includes a retained residual tensor and a retained scale tensor;

[0016] Convert the precision of the residual tensor in the pair of retained elements to a preset precision, and convert the scale tensor in the pair of retained elements into index information with a preset number of bits;

[0017] Recombine the pair of retained elements in the data buffer to output a masked residual tensor, a masked scale tensor, and valid element information.

[0018] In some feasible embodiments, the recombination conversion table includes a first recombination conversion table and a second recombination conversion table; wherein, the first recombination conversion table is used for residual encoding, and the second recombination conversion table is used for super-tensor encoding;

[0019] When the coding rate is changed, the first recombination conversion table remains unchanged, and the second recombination conversion table needs to be switched to the second recombination conversion table corresponding to the coding rate.

[0020] In some feasible embodiments, when performing encoding, the residual encoding module is specifically configured as:

[0021] Read the first recombination conversion table and the recombination status table in the on-chip cache unit, and obtain the quantity of the input data, calculate a first value and a first remainder, where the first value is the result of rounding down the input data volume to the nearest multiple of 4, and the first remainder is the remainder generated after performing a division operation on the input data volume by 4;

[0022] If the first remainder is 0, perform external residual encoding and internal residual encoding on the data of the first value. The external residual encoding is used to perform clipping processing on the elements in the masked residual tensor so that the range of the elements is within a first preset range. The internal residual encoding is used to encode the elements after the clipping processing to generate a residual bitstream, and the number of bits of the residual bitstream is the number of bits within a preset range;

[0023] If the first remainder is non - zero, perform external residual encoding and internal residual encoding in a first preset order based on the first recombination conversion table and the recombination status table to generate a residual bitstream. The first preset order includes a first order and a second order. The first order is the order corresponding to the data with the number of the first remainder, and the second order is the order corresponding to the data with the number of the first value.

[0024] In some feasible embodiments, when the super - tensor encoding module performs encoding, it is specifically configured to:

[0025] Read the second recombination conversion table in the on - chip cache unit and perform per - channel encoding on the super - tensor based on the second recombination conversion table;

[0026] Obtain the number of the super - tensors, and calculate a second value and a second remainder. The second value is the result of rounding down the number of the super - tensors to the nearest multiple of 4, and the second remainder is the remainder generated after performing a division operation on the number of the super - tensors by 4;

[0027] If the second remainder is 0, perform external residual encoding and internal residual encoding on the data of the second value. The external residual encoding is used to perform clipping processing on the elements of the super - tensor so that the range of the elements of the super - tensor is within a second preset range. The internal residual encoding is used to encode the elements of the super - tensor after the clipping processing to generate a super - tensor bitstream;

[0028] If the second remainder is non - zero, perform external residual encoding and internal residual encoding in a second preset order to generate a super - tensor bitstream. The second preset order includes a third order and a fourth order. The third order is the order corresponding to the data with the number of the second remainder, and the second order is the order corresponding to the data with the number of the second value.

[0029] In some feasible embodiments, both the residual encoding module and the super - tensor encoding module include a storage structure. The storage structure includes a primary storage structure, a secondary storage structure, and a tertiary storage structure;

[0030] The width of the first-level storage structure is the first width. The second-level storage structure includes a first number of first storage units, and the width of the first storage unit is the second width. The third-level storage structure includes a second number of second storage units, and the width of the second storage unit is the third width;

[0031] The second width is one-half of the first width, and the third width is the product of the first number and the second width.

[0032] In some feasible embodiments,

[0033] The first-level storage structure of the residual coding module and the super-tensor coding module is used to receive the residual code stream and the super-tensor code stream, shift the code rate in the first direction by the current number of bits, and perform a bitwise OR operation with the data already stored in the slots of the first storage structure to store the residual code stream;

[0034] After a second preset number of residual codings and / or super-tensor codings are completed, obtain the status of the slots;

[0035] If the status of the slots is a half-full state, the second-level storage structure of the residual coding module and the super-tensor coding module is used to retrieve and store the data that meets the preset conditions written in the first storage structure, and the preset condition is the data with a high write priority;

[0036] Obtain the storage status of the second storage structure;

[0037] If the storage status is a full state, the third-level storage structure of the residual coding module and the super-tensor coding module reads and stores the residual code stream and the super-tensor code stream stored in the second-level storage structure.

[0038] In some feasible embodiments, it further includes: a bus controller and an off-chip memory. The bus controller communicates with the host through a general protocol and is used to receive the configuration information of the host. The configuration information includes a start signal, a mode code, address information, and input data scale information;

[0039] When the bus controller receives the start signal and the mode code is the mask unit mode, it sends a first start signal to the mask unit to control the mask unit to perform block processing on the input data and perform a mask operation on the input data; and, based on the address information, sends a first control signal to the mask unit, the data loading unit, and the write-back unit to control the data loading unit and the write-back unit to complete the data interaction between the on-chip cache unit and the off-chip memory;

[0040] When the bus controller receives a first control signal and the mode code is the loading lookup table mode of the data loading unit, it sends a second control signal to the data loading unit based on the address information to control the data loading unit to load the conversion table and the status table and store them in the on-chip cache unit;

[0041] When the bus controller receives a start signal and the mode code is the residual encoding module mode, it sends a second start signal to the residual encoding module to control the residual encoding module to perform block processing based on the output length of the mask unit and perform an encoding operation; meanwhile, based on the address information, it sends a second control signal to the residual encoding module, the data loading unit, and the write-back unit to control the data loading unit and the write-back unit to complete the data interaction between the on-chip cache unit and the off-chip memory;

[0042] When the mode code is the super-tensor encoding module mode, the bus controller further sends a third start signal to the super-tensor encoding module to control the super-tensor encoding module to perform block processing based on the number of channels and then perform an encoding operation; and based on the address information, it sends a third control signal to the super-tensor encoding module, the data loading unit, and the write-back unit to control the data loading unit and the write-back unit to complete the data interaction between the on-chip cache unit and the off-chip memory.

[0043] In a second aspect, the present application further provides a JPEG-AI-based entropy encoding method, including:

[0044] Obtain input data, where the input data includes a residual tensor, a scale tensor, and a super-tensor;

[0045] Perform a masking operation on the input data to output a masked residual tensor, a masked scale tensor, and valid element information;

[0046] Recombine the conversion table and the status table to output a recombined conversion table and a recombined status table;

[0047] Encode the masked residual tensor based on the masked scale tensor, the recombined conversion table, the recombined status table, and the valid element information to output a residual bitstream; and encode the super-tensor based on the recombined conversion table to output a super-tensor bitstream.

[0048] As can be seen from the above technical solutions, the present application provides an entropy coding hardware accelerator based on JPEG-AI and an entropy coding method, which can complete image compression with arbitrary input resolution and arbitrary bit rate. The accelerator includes a mask unit, a data loading unit, an on-chip cache unit, an encoding unit, and a data write-back unit. Among them, the encoding unit includes a residual encoding module and a super-tensor encoding module. The mask unit is used to perform mask operations on the residual tensor and the scale tensor, and output a masked residual tensor, a masked scale tensor, and valid element information. The on-chip cache unit is used to store the reorganized conversion table, the reorganized status table, the input data and the output data of the encoding unit loaded by the data loading unit. The residual encoding module is used to perform encoding on the masked residual tensor based on the masked scale tensor, the reorganized conversion table, the reorganized status table, and the valid element information, so as to output a residual bitstream. The super-tensor encoding module is used to perform encoding on the super-tensor based on the reorganized conversion table, so as to output a super-tensor bitstream. The accelerator reorganizes the conversion table and the status table and stores them in the on-chip cache unit, which can reduce the calculation and access storage time. At the same time, if the bit rate is not changed, the table corresponding to the current bit rate stored on the chip can be directly used when compressing other pictures, thus avoiding re-accessing the memory and improving the execution efficiency of entropy coding and the efficiency of image coding. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0050] Figure 1 Schematic structural diagram of the entropy coding hardware accelerator based on JPEG-AI provided by the embodiment of the present application;

[0051] Figure 2 Schematic structural diagram of the data buffer of the mask unit provided by the embodiment of the present application;

[0052] Figure 3 Schematic diagram of the build_index operation process provided by the embodiment of the present application;

[0053] Figure 4 Schematic diagram of the in-group encoding process provided by the embodiment of the present application;

[0054] Figure 5 Schematic diagram of the three-level storage structure provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] Embodiments will be described in detail below, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following examples do not represent all embodiments consistent with the present application. They are merely examples of systems and methods consistent with some aspects of the present application detailed in the claims.

[0056] JPEG AI uses a deep learning model to analyze and encode images. The deep learning model is trained with a large amount of image data and can learn the statistical features and semantic information of the images, thereby achieving high-quality and high-bitrate image compression.

[0057] Entropy coding is an integral part of the JPEG AI image coding standard and is a lossless data compression technique. Entropy coding assigns different lengths of codes according to the probability distribution of the data. Specifically, data with a high occurrence probability is assigned a shorter code, while data with a low occurrence probability is assigned a longer code. When encoding the entire data sequence, the average coding length will be relatively short, thereby achieving the purpose of compressing the data.

[0058] During the JPEG AI encoding process, the input image is first subjected to neural network calculations to obtain a latent representation, and then entropy coding is applied to this latent representation to obtain a bitstream. By performing entropy coding on the latent representation, the data volume can be significantly reduced, and the costs of image storage and transmission can be lowered.

[0059] During the JPEG AI encoding process, the input image will undergo neural network calculations. The neural network hardware accelerator can efficiently process operations such as convolution and activation, and quickly convert the input image into a latent representation. This latent representation can be understood as an intermediate data form after feature extraction and transformation. Then, entropy coding processing is performed on the obtained latent representation, and the latent representation is further converted into a bitstream through entropy coding to achieve data compression and storage transmission.

[0060] Thus, it can be seen that neural network calculation is a prerequisite step for entropy coding and provides the data to be encoded. It can be understood that if the neural network hardware accelerator has strong performance and can quickly complete neural network calculations, the speed of obtaining the latent representation will be accelerated. The entropy coding module can obtain the data to be encoded faster, thereby starting the encoding operation, and the overall compression process time will be shortened. On the contrary, if the neural network calculation speed is slow, the entropy coding module needs a long time to obtain the data, resulting in low efficiency in the encoding process.

[0061] Even if a neural network hardware accelerator can efficiently complete computational tasks, if the entropy encoding process is inefficient and consumes high power, it will also affect the performance of image encoding. Since entropy encoding is the step of converting the latent representation into a bitstream, its efficiency directly determines the speed of bitstream generation. If the entropy encoding efficiency is not high, it will cause the output speed of the encoding system to slow down and fail to meet the real-time requirements.

[0062] The entropy encoding algorithm specified by the JPEG AI standard can only be executed by the CPU. Due to the low efficiency and high power consumption of the CPU when processing entropy encoding tasks, the bitstream generation speed is slow, resulting in low image encoding efficiency.

[0063] To solve the problem of low image encoding efficiency, some embodiments of this application provide an entropy encoding hardware accelerator based on JPEG-AI to complete image compression with arbitrary input resolutions and arbitrary bitrates. The accelerator can reduce the calculation and access storage time by reorganizing the conversion table and the state table and storing them in the on-chip cache unit. At the same time, if the bitrate is not changed, the table corresponding to the current bitrate stored on the chip can be directly used when compressing other pictures, thus avoiding re-accessing the memory and improving the execution efficiency of entropy encoding and the efficiency of image encoding.

[0064] As Figure 1 shown, the accelerator provided by some embodiments of this application can support input images of various resolutions and various compression bitrates at the same time, and can meet the entropy encoding requirements of the JPEG AI standard. The accelerator includes a mask unit, a data loading unit, an on-chip cache unit, an encoding unit, and a data write-back unit. Among them, the encoding unit includes a residual encoding module and a super-tensor encoding module.

[0065] The mask unit is used to perform mask operations on the residual tensor and the scale tensor. The residual encoding module is used to perform residual encoding, and the super-tensor encoding module is used to perform super-tensor encoding. The on-chip cache unit is a collection of multiple pseudo-dual-port SRAMs (Static Random-Access Memory).

[0066] The mask unit is used to perform mask operations on the input data, and the input data includes a residual tensor (res) and a scale tensor (scale). Among them, the residual tensor is the main encoding information, which contains the remaining information after the image is processed and is important for image reconstruction. The scale tensor is auxiliary encoding information, and its dimension is the same as that of the residual tensor, and the data corresponds one by one, and is used in the mask unit to determine whether to retain the corresponding res data.

[0067] In some embodiments, when performing mask operations, the mask unit is specifically configured as:

[0068] If an element of the scale tensor is greater than a preset value, mark the element as a true element, obtain a corresponding element in the residual tensor based on the element to obtain a retained element pair, and write the retained element pair into a data buffer. The retained element pair includes a retained residual tensor and a retained scale tensor, and a data buffer is provided in the mask unit;

[0069] Convert the precision of the residual tensor in the retained element pair to a preset precision, and convert the scale tensor in the retained element pair into index information with a preset number of bits;

[0070] The retained element pairs are reorganized in the data buffer to output a masked residual tensor, a masked scale tensor, and valid element information.

[0071] When performing a masking operation, first determine whether to retain the corresponding res and scale data pair according to the value of scale. Among them, the preset value is 382. Determine whether the element of the scale data is greater than 382. If a certain scale element is greater than 382, it means that this pair of res and scale data meets the retention condition, that is, "the mask is true", and it is a retained element pair. If the scale element is less than or equal to 382, then "the mask is false", and this data pair will be discarded.

[0072] As Figure 2 shown, after preprocessing the retained res and scale data, write them into a data buffer with a data length of 32×2 in ascending order in a loop. The data buffer is divided into a lower half area and a higher half area, that is, a first data storage unit and a second data storage unit. A masking operation for a pair of data is performed in each clock cycle, and the higher and lower half areas are used alternately. If one of the half areas is full, the data in it can be written into the SRAM, and the masking operation continues to write data in the other half area without stagnation, so as to improve the continuity and efficiency of data processing.

[0073] After it is determined that "the mask is true", the corresponding res and scale data pair, that is, the retained element pair, will enter the subsequent processing link to perform data preprocessing, that is, perform format conversion on the retained element pair, and preprocess the res and scale data respectively. Among them, the res data will be converted from fp32 to a preset precision by a precision conversion unit, and the preset precision is int16.

[0074] The scale data will perform a build_index operation, as Figure 3As shown, the build_index operation is a conversion operation for scale data, converting from 32-bit scale to index information with a preset number of bits. The preset number of bits is 8 bits. Then, the processed data is written into the output buffer. Specifically, adding 64 to the scale data can perform a certain offset adjustment on the data to make the data distributed within a suitable range. Then, shift the result after adding 64 to the right by 7 bits. The right shift operation is equivalent to dividing the data by 128, which plays a role in scaling the data and further adjusting the data range. Then, limit the data after shifting right by 7 bits within the range of 0 to 34. If the data is less than 0, take 0. If it is greater than 34, take 34, which can make the data within a reasonable range. Take the lower 8 bits from the limited data, and finally obtain 8-bit index information, which can reduce the storage space occupied by the data and is also convenient for subsequent encoding and other operations.

[0075] When storing the same amount of scale data, the required storage space is significantly reduced. In the image coding scenario, especially when processing a large amount of image data, the storage resources are limited. This reduction in data bit width can effectively save storage space and improve storage efficiency.

[0076] Moreover, during the data processing, compared with the 32-bit scale data, the 8-bit index information requires less bandwidth for transmission and can be transmitted faster between different hardware modules, such as the MASK unit, ENCODE_R unit, etc., reducing the data transmission delay and thus improving the processing speed of the entire encoding.

[0077] The converted 8-bit index information can be used as an index to facilitate quick searching and mapping in the recombination conversion table and the recombination status table. During the encoding process, some lookup tables may be used for data conversion and processing. The 8-bit index can more efficiently locate the corresponding position in the lookup table, thus accelerating the encoding speed.

[0078] The remaining data, that is, the residual tensor after conversion accuracy and the scale tensor converted to the preset number of bits, will be recombined in a compact form to form the tensors res_mask (masked residual tensor) and scale_mask (masked scale tensor) that are continuously stored in memory. Data recombination can improve the efficiency of data access and processing. At the same time, the mask unit will count the number mask_true_num (32 bits) of the masked true values (i.e., the remaining data), which is the valid data information, and retain the valid data information in the register for use by the residual coding module. When performing the encoding operation, the residual coding module needs mask_true_num to adjust the encoding strategy, allocate resources, or control the encoding process.

[0079] In the standard algorithm, when the residual encoding module encodes the masked residual tensor res_mask and the scale tensor scale_mask, it needs to calculate the relevant lookup table online. This online calculation method consumes a large amount of computing resources and time. Especially when processing a large number of images, it will lead to low encoding efficiency and increase the power consumption of the hardware. In this embodiment, the tables that the residual encoding module in the standard algorithm needs to calculate online are solidified outside the chip to obtain the fixed transr_table (the first recombination transformation table) and state_table (the recombination state table), and the in-memory computing method is adopted to store the tables in the on-chip SRAM. After the accelerator is powered on, it only needs to be loaded once and can be reused.

[0080] Similarly, when the super-tensor encoding module encodes the super-tensor z, it needs to calculate the relevant lookup table online. In this embodiment, the tables that the super-tensor encoding module in the standard algorithm needs to calculate online are solidified outside the chip to obtain the fixed transz_table (the second recombination transformation table), and the in-memory computing method is adopted to store the tables in the on-chip SRAM. After the accelerator is powered on, it only needs to be loaded once and can be reused by tasks with the same compression bit rate.

[0081] By using the in-memory computing method to store the solidified transr_table, state_table, and transz_table in the on-chip SRAM, the transmission overhead between the storage unit and the computing unit can be reduced, and the data processing speed can be improved. The on-chip SRAM has the characteristics of high-speed reading and writing and can quickly provide the lookup table data required for encoding.

[0082] The data loading unit is used to receive the loading instruction and load the recombination transformation table, recombination state table, residual tensor, scale tensor, and super-tensor in the off-chip memory into the on-chip cache unit. The on-chip cache unit is used to store the recombination transformation table, recombination state table, and the input data and output data of the encoding unit. Among them, the input data includes the residual tensor, scale tensor, and super-tensor, and the output data includes the residual bitstream and the super-tensor bitstream.

[0083] The data width of transr_table is 32, and the dimension is (35, 256). The data width of state_table is 8, and the dimension is (35, 256). In this embodiment, the dimension of transr_table is recombined into (140, 64), so that the number of bits in each row of the two lookup tables is the preset number of bits, that is, 2048 bits. After this mode is started, the residual encoding module issues the continuously transmitted commands to the data loading unit. The loaded data of the first 140 rows is written into transr_sram, that is, the first cache table storage module, and the loaded data of the last 35 rows is written into state_y_sram, that is, the state table storage module.

[0084] In addition, bound_table is also required to perform clipping processing on the masked residual tensor. bound_table is an array with a length of 35. The residual encoding module queries bound_table through the 8-bit index information in scale_mask to perform clipping processing on res_mask. When the residual encoding module performs residual encoding, based on the conversion mapping relationship in transr_table, the data in the clipped res_mask is converted into a coded bitstream form; at the same time, according to the state transition information recorded in state_table, the encoding state value is adjusted to complete continuous encoding.

[0085] The residual encoding module performs encoding based on the masked residual tensor, the recombination conversion table, the recombination state table, and the valid element information to output a residual bitstream. In some embodiments, when the residual encoding module performs encoding, it is specifically configured to:

[0086] Obtain the quantity of the input data, and calculate a first value and a first remainder. The first value is the result of rounding down the input data volume to the nearest multiple of 4, and the first remainder is the remainder generated after performing a division operation on the input data volume by 4;

[0087] If the first remainder is 0, perform external residual encoding and internal residual encoding on the data of the first value. The external residual encoding is used to perform clipping processing on the elements in the masked residual tensor so that the range of the elements is within a first preset range, and the internal residual encoding is used to perform encoding on the clipped elements to generate a residual bitstream. The number of bits of the residual bitstream is the number of bits within the preset range;

[0088] If the first remainder is non-zero, perform external residual encoding and internal residual encoding in a first preset order based on the first recombination conversion table and the recombination state table to generate a residual bitstream. The first preset order includes a first order and a second order. The first order is the order corresponding to the data of the first remainder quantity, and the second order is the order corresponding to the data of the first value quantity.

[0089] When the residual encoding module performs encoding, each data r in res_mask needs to go through outbound_r encoding (external residual encoding) and inbound_r encoding (internal residual encoding). During encoding, first take the remainder of the input number divided by 4. If the obtained remainder is n, perform outbound_r encoding and inbound_r encoding on the first n data in the input data in sequence. At this time, the remaining input quantity is a multiple of 4, and then group the remaining input in groups of 4, and perform encoding on each group in sequence. The encoding process within the group is as followsFigure 4 as shown

[0090] outbound_r limits the range of r to [-bound, bound–1], where bound = bound_table[index], bound_table is an array of length 35, reserved as a constant in the hardware design, and index is an element in scale_mask. If the absolute value of r is less than bound, r remains unchanged; if the absolute value of r is greater than or equal to bound, then r = -bound, and the encoded value valueCoded is written into the bitstream. The number of bits of valueCoded depends on whether it is greater than or equal to 8. If it is greater than or equal to 8, it is 14 bits, otherwise it is 4 bits.

[0091] valueCoded is calculated according to the following formula:

[0092] sign = r >> 15;

[0093] valueCoded = ((r + (sign ^ (-bound))) << 1) ^ sign;

[0094] The function of inbound_r is to encode the clipped r into resCoded, and the number of bits of resCoded, resBits, is 0 to 8. In some embodiments, the internal residual coding is performed through the first state variable (state_1) and the second state variable (state_2). The initial values of the state variables state_1 and state_2 are 0. For all r in res_mask, the encoded values of the r in odd order come from state_1, and the encoded values of the r in even order come from state_2. resCoded takes the lower resBits bits of state1 or state_2.

[0095] The calculation method of resBits in each inbound_r encoding and the update method of state_1 / 2 after encoding are as follows:

[0096] transr = transr_table[index][r];

[0097] resBits = (state_1 / 2 + (transr >> 16)) >> 8;

[0098] state_1 / 2 = ((state_1 / 2 | 256) >> resBits) + transr;

[0099] First, obtain transr from the lookup table transr_table according to index and r. Then, calculate resBits based on transr and the current state_1 or state_2, which is the number of bits of the encoded value resCoded. Next, update state_1 or state_2 according to resBits for use in the next encoding. In this way, alternately use state_1 and state_2 to process r in odd and even orders to complete the encoding process for all r in res_mask.

[0100] In some embodiments, the super-tensor encoding module performs per-channel encoding on the super-tensor z based on the second recombination transformation table; the data bit-width of the second recombination transformation table transz_table is 32, and the dimension is (192, 64). The second recombination transformation tables corresponding to different code rates are pre-calculated outside the chip, and then the data loading unit writes the table corresponding to the specified code rate into transz_sram, that is, the second cache table storage module. For tasks with the same encoding code rate, different pictures can reuse the same table without reloading.

[0101] Both the masking operation and the encoding operation of res depend on scale. Since scale is obtained by neural network calculation of the super-tensor z, it is necessary to perform super-tensor encoding on z to retain scale information for decoding. The dimension of z is (192, size_z), and the data type is uint8. However, since the value of the element value_z is clipped to [0, 63], there are only 6 valid bits.

[0102] The super-tensor encoding module performs per-channel encoding on 192 channels of the super-tensor z. transz_table has 192 rows and 64 columns, and each row corresponds to each channel of z. Therefore, before starting the encoding of the nth channel, it is necessary to read out 2048 bits of data (including 64 32-bit numbers) in the nth row of transz_sram to assist in the encoding of z.

[0103] In some embodiments, when performing encoding, the super-tensor encoding module is specifically configured to:

[0104] Read the second recombination transformation table of the second cache table storage module and perform per-channel encoding on the super-tensor based on the second recombination transformation table;

[0105] Obtain the number of the super-tensors, and calculate a second value and a second remainder. The second value is the result of rounding down the number of super-tensors to the nearest multiple of 4, and the second remainder is the remainder generated after performing a division operation on the number of super-tensors by 4;

[0106] If the second remainder is 0, perform external residual coding and internal residual coding on the data of the second value. The external residual coding is used to perform clipping processing on the elements of the super-tensor so that the range of the elements of the super-tensor is within a second preset range, and the internal residual coding is used to perform coding on the elements of the super-tensor after the clipping processing to generate a super-tensor bitstream;

[0107] If the second remainder is non-zero, perform external residual coding and internal residual coding in a second preset order to generate a super-tensor bitstream. The second preset order includes a third order and a fourth order. The third order is the order corresponding to the data of the second remainder quantity, and the second order is the order corresponding to the data of the second value quantity.

[0108] For the size_z data within a single channel of the super-tensor z, adopt a grouping method of taking the remainder of 4 in the same way as the residual coding unit. First, perform outbound_z coding and inbound_z coding on the first n numbers (the remainder of size_z divided by 4) successively, and then group the remaining inputs by 4, and perform continuous outbound_z and continuous inbound_z coding within the group.

[0109] outbound_z is used to query whether the transz corresponding to value_z is 0. transz is a value in the lookup table transz_table, and is a numerical value obtained from the transz_table according to the element value_z in the super-tensor z and the current channel index channel_index during the outbound_z coding. If transz is 0, write the 6-bit value_z into the bitstream and re-assign value_z to 63.

[0110] inbound_z further encodes value_z into zCoded, and the number of bits zBits of zCoded is 0 to 8. inbound_z also relies on two state variables state_1 and state_2 that are initially 0, which are used for odd-numbered coding and even-numbered coding respectively. The calculation method of the output length zBits during each inbound_z coding, and the update method of state_1 and state_2 after coding are as follows:

[0111] transz = transz_table[channel_index][value_z];

[0112] zBits = (state_1 / 2 + (transz >> 16)) >> 8;

[0113] state_1 / 2 = ((state_1 / 2 | 256) >> zBits) + transz;

[0114] In summary, the super-tensor encoding module performs encoding channel by channel on 192 channels of the super-tensor z. Before encoding each channel, the corresponding row data is read from transz_sram. For the data within the channel, the remainder part is processed first, and then the remaining data is grouped for encoding. During the encoding process, the outbound_z encoding determines whether to write value_z into the code stream and update value_z according to the value of transz. The inbound_z encoding uses the state variables state_1 and state_2 to further encode value_z into zCoded. Such an encoding method can effectively encode the relevant information z of scale into the code stream, and at the same time, combined with the encoding of res, a complete entropy encoding process is achieved.

[0115] The encoding length of each data is uncertain, and the output of continuous encoding needs to be integrated into a regular length of 512 bits (the width of the SRAM for storing the code stream). To solve the problems of uncertain encoding length and difficult memory width alignment, in some embodiments, both the residual encoding module and the super-tensor encoding module include a storage structure, and the storage structure includes a primary storage structure, a secondary storage structure, and a tertiary storage structure;

[0116] The width of the primary storage structure is the first width, the secondary storage structure includes a first number of first storage units, and the width of the first storage unit is the second width. The tertiary storage structure includes a second number of second storage units, and the width of the second storage unit is the third width;

[0117] The second width is half of the first width, and the third width is the product of the first number and the second width.

[0118] For the storage structure of the residual encoding module, in some embodiments, the primary storage structure of the residual encoding module is used to receive the residual code stream, and move the code stream in the first direction by the current number of bits, and perform a bitwise OR operation with the data already stored in the slot of the first storage structure to store the residual code stream;

[0119] For the storage structure of the super-tensor encoding module, in some embodiments, the primary storage structure of the super-tensor encoding module is used to receive the super-tensor code stream, and move in the first direction by the current number of bits, and perform a bitwise OR operation with the data already stored in the slot of the first storage structure to store the super-tensor code stream;

[0120] After the second preset number of residual encodings or super-tensor encodings are completed, obtain the slot status;

[0121] If the slot status is in the half-full state, the secondary storage structure of the encoding module is used to retrieve and store the data meeting the preset conditions written in the first storage structure, where the preset conditions are data with a high write priority;

[0122] Obtain the storage status of the second storage structure;

[0123] If the storage status is in the full state, the tertiary storage structure of the residual and super-tensor encoding module reads and stores the residual code stream and the super-tensor code stream stored in the secondary storage structure.

[0124] In the residual encoding module and the super-tensor encoding module, there is the same tertiary storage structure, as Figure 5 shown, including a primary storage structure bitsteam_slot, a secondary storage structure bitstream_group, and a tertiary storage structure bitstream_sram.

[0125] bitsteam_slot is a register with a width of 128 bits, which is used to directly store the encoding output. The number of valid bits therein is recorded by current_bits. Each time a new encoding output is generated, it will be shifted left by current_bits and then perform a bitwise OR operation with the data in the slot, so as to be written into the slot. The encoding module will regularly check whether the slot is half-full. If it is half-full, the first 64 bits written will be loaded into the secondary storage bitstream_group.

[0126] Exemplarily, when new encoded code stream data is generated, the new encoded data needs to be shifted left by current_bits bits to set the storage position of the new encoded data in bitsteam_slot, so that the new data can be stored immediately following the existing valid data in bitsteam_slot. For example, assume that the value of current_bits is 32 and the new encoded data is 00110011, then 00110011 will be shifted left by 32 bits. After the left shift operation, perform a bitwise OR operation on the new encoded data and the existing data in bitsteam_slot. The rule of the bitwise OR operation is that as long as one of the corresponding two binary bits is 1, the result bit is 1, otherwise it is 0. Through the bitwise OR operation, the new encoded data and the existing data in bitsteam_slot can be combined together to achieve the writing of new data.

[0127] The bitstream_group is a register group with a size of 8×64 bits. When the bitsteam_slot is half full, the first-written 64-bit data, that is, the data with higher priority, is loaded into it. After the bitstream_group is loaded with data 8 times, that is, 512 bits, the data is written to the three-level memory bitstream_sram.

[0128] The bitstream_sram is a SRAM with a width of 512 bits, which is used to store the 512-bit data written from the bitstream_group. In some embodiments, the accelerator further includes a write-back unit and an off-chip memory. After encoding is completed, the encoding unit sends a control signal to the write-back unit, and the write-back unit writes the code stream back to the off-chip memory.

[0129] By setting the three-level memory structure, the output encoding with an uncertain length can be integrated into a regular 512-bit length under the conditions of lower power consumption and area, which is convenient for subsequent data storage and processing.

[0130] For the off-chip DDR, when the encoding unit and the mask unit need to read the input data of the off-chip DDR, they need to send a control signal to the data loading unit first. After waiting for the data loading unit to load the data from the DDR to the on-chip cache SRAM, they then read from the SRAM. Similarly, when the encoding unit and the mask unit need to write the output data back to the DDR, they both need to write the data into the SRAM first, and then send a control signal to the write-back unit so that the write-back unit writes the data in the SRAM back to the DDR.

[0131] In some embodiments, the accelerator further includes a bus controller CONTROL. The bus controller communicates with the host through the general protocol AXI4-Lite, and is used to receive the configuration information of the host. The configuration information includes a start signal, a mode code, address information, input size information, and at the same time, information such as an end signal and the length of the output code stream can also be returned to the host.

[0132] Exemplarily, the following is the configuration information sent by the host to the accelerator:

[0133] Start signal: valid, with a bit width of 1 bit. After being pulled high, the accelerator starts to work;

[0134] Mode code: mode_id, 4 bits. 4’b0000 represents the mask mode, 4’b0001 represents the residual encoding mode, 4’b0010 represents the super-tensor encoding mode, 4’b0011 represents the residual loading mode, and 4’b0100 represents the super-tensor loading mode;

[0135] Input addresses ddr_raddr_0 and ddr_raddr_1: 32 bits, used to represent the storage addresses of the input data in the off-chip memory;

[0136] Output addresses ddr_waddr_0 and ddr_waddr_1: 32 bits, used to represent the storage addresses of the output data in the off-chip memory;

[0137] Input length: input_length, 20 bits, characterizing the length of the input data for this task; Channel size: size_per_channel, 16 bits, the dimension of the super-tensor z is set to (192, size_per_channel);

[0138] Number of channels per block: channel_per_block, 8 bits, used to block the z encoding work, with a maximum of 192.

[0139] When the bus controller receives a start signal and the mode code is the mask unit mode, it sends a first start signal to the mask unit to control the mask unit to perform block processing based on the input data and perform a mask operation on the input data; and, based on the address information, sends a first control signal to the mask unit, data loading unit, and write-back unit to control the data loading unit and write-back unit to complete the data interaction between the on-chip cache unit and the off-chip memory.

[0140] Exemplarily, configure the mask unit. CONTROL receives control signals from the host computer, mode_id is set to 4'b0000, ddr_raddr_0 and ddr_raddr_1 are respectively set to the addresses of the input res and scale, ddr_waddr_0 and ddr_waddr_1 are respectively set to the addresses of the output res_mask and scale_mask, and input_length is set to the length of the input data. Then perform the mask operation. CONTROL outputs a valid signal to the mask unit. The mask unit blocks the work according to the input length, and each block performs a mask operation on at most 32768 input data. The mask unit can send control signals to the data loading unit and write-back unit to load the input data from the DDR to the on-chip SRAM and write the output data from the on-chip SRAM back to the DDR.

[0141] When the bus controller receives a first control signal and the mode code is the loading lookup table mode of the data loading unit, it sends a second control signal to the data loading unit based on the address information to control the data loading unit to load the conversion table and status table and store them in the on-chip cache unit.

[0142] Exemplarily, when CONTROL receives a control signal and mode_id is set to 4'b0011, it will transmit control information to the data loading unit to load the first recombination conversion table and the recombination status table.

[0143] Again exemplarily, when CONTROL receives a control signal and mode_id is set to 4'b0100, it will transmit control information to the data loading unit to load the second recombination conversion table.

[0144] The bus controller is used to send a second start signal to the residual encoding module when receiving a start signal and the mode code is the residual encoding module mode, so as to control the residual encoding module to perform block processing based on the output length of the mask unit and perform an encoding operation; meanwhile, based on the address information, send a second control signal to the residual encoding module, the data loading unit and the write-back unit to control the data loading unit and the write-back unit to complete the data interaction between the on-chip cache unit and the off-chip memory.

[0145] Exemplarily, the residual encoding module is configured, CONTROL receives a control signal from the host computer, mode_id is set to 4'b0001, ddr_raddr_0 and ddr_raddr_1 are respectively set to the addresses of the input res_mask and scal_mask, and ddr_waddr_0 is set to the address of the output bitstream_r. Then perform the ENCODE_R operation, CONTROL outputs a valid signal to the residual encoding module, and the residual encoding module divides the work according to the output length of the mask unit. Each block performs an encoding operation on up to 32,768 input data at most. The residual encoding module can send control signals to the data loading unit and the write-back unit to load the input data from the DDR to the on-chip SRAM and write the output data from the on-chip SRAM back to the DDR.

[0146] The bus controller is further used to send a third start signal to the super-tensor encoding module when the mode code is the super-tensor encoding module mode, control the super-tensor encoding module to perform block processing based on the number of channels, and then perform an encoding operation; and based on the address information, send a third control signal to the super-tensor encoding module, the data loading unit and the write-back unit to control the data loading unit and the write-back unit to complete the data interaction between the on-chip cache unit and the off-chip memory.

[0147] Exemplarily, configure the super-tensor encoding module. CONTROL receives control signals from the host computer. mode_id is set to 4'b0010, ddr_raddr_0 is set to the address of the input super-tensor z, ddr_waddr_0 is set to the address of the output bitstream_z, size_per_channel is set to the dimension of each channel, and channel_per_block is set to the number of channels per block. Then perform the ENCODE_Z operation. CONTROL outputs a valid signal to the super-tensor encoding module. The super-tensor encoding module divides the work into blocks according to channel_per_block. The super-tensor encoding module can send control signals to the data loading unit and the write-back unit, load the input data from the DDR to the on-chip SRAM, and write the output data from the on-chip SRAM back to the DDR.

[0148] By setting the block logic in the mask unit, the residual encoding module, and the super-tensor encoding module, it can be ensured that large-size images can be encoded under the condition of limited on-chip cache.

[0149] Under the existing framework, the accelerator provided in this embodiment only needs to give the operating mode, input size, and data address to implement the complete entropy encoding process, which places a relatively small burden on programmers.

[0150] Based on the entropy encoding hardware accelerator based on JPEG-AI provided in the above embodiment, some embodiments of the present application also provide an entropy encoding method based on JPEG-AI, including:

[0151] Obtain input data, where the input data includes a residual tensor, a scale tensor, and a super-tensor;

[0152] Perform a masking operation on the input data to output a masked residual tensor, a masked scale tensor, and valid element information;

[0153] Recombine the conversion table and the state table to output a recombined conversion table and a recombined state table;

[0154] Based on the masked scale tensor, the recombined conversion table, the recombined state table, and the valid element information, encode the masked residual tensor to output a residual bitstream; and based on the recombined conversion table, encode the super-tensor to output a super-tensor bitstream.

[0155] The present application provides an entropy coding hardware accelerator and an entropy coding method based on JPEG-AI, which can complete image compression with any input resolution and any bit rate. The accelerator includes a masking unit, a data loading unit, an on-chip cache unit, an encoding unit, and a data write-back unit. The encoding unit includes a residual encoding module and a super-tensor encoding module. The masking unit performs a masking operation on the residual tensor and the scale tensor, and outputs a masked residual tensor, a masked scale tensor, and valid element information. The on-chip cache unit stores the reorganized transformation table, the reorganized status table, the input data and the output data of the encoding unit output by the data loading unit. The residual encoding module encodes the masked residual tensor based on the masked scale tensor, the reorganized transformation table, and the reorganized status table, and outputs a residual bitstream. The super-tensor encoding module loads a super-tensor and encodes the super-tensor based on the reorganized transformation table to output a super-tensor bitstream. At the same time, if the bit rate is not changed, the table corresponding to the current bit rate stored on the chip can be directly used when compressing other pictures, thereby avoiding re-accessing the memory and improving the execution efficiency of entropy coding and the efficiency of image coding. If the bit rate is changed, the corresponding partial table is reloaded to ensure the accuracy and effectiveness of the encoding.

[0156] For the similarities between the embodiments provided in the present application, reference may be made to each other. The specific embodiments provided above are only several examples under the general concept of the present application and do not constitute a limitation on the protection scope of the present application. For those skilled in the art, any other embodiments extended based on the solution of the present application without creative efforts fall within the protection scope of the present application.

Claims

1. An entropy coding hardware accelerator based on JPEG-AI, characterized in that, Comprising: A mask unit for performing a mask operation on input data to output a mask residual tensor, a mask scale tensor, and valid element information, where the input data includes a residual tensor and a scale tensor; A data loading unit for loading a recombined conversion table, a recombined status table, a residual tensor, a scale tensor, and a hyper tensor from an off-chip memory into an on-chip cache unit, where the recombined conversion table is the converted table after recombination, and the recombined status table is the status table after recombination; An on-chip cache unit for storing the recombined conversion table and the recombined status table; An encoding unit including a residual encoding module and a hyper tensor encoding module, where the residual encoding module encodes the mask residual tensor based on the mask scale tensor, the recombined conversion table, the recombined status table, and the valid element information to output a residual bitstream; The hyper tensor encoding module encodes the hyper tensor based on the recombined conversion table to output a hyper tensor bitstream; A data write-back unit for reading the mask residual tensor, the mask scale tensor, the residual bitstream, and the hyper tensor bitstream stored in the on-chip cache unit and writing them back to the off-chip memory.

2. The entropy encoding hardware accelerator based on JPEG-AI according to claim 1, wherein A data buffer is provided in the mask unit, and the data buffer includes a first data storage unit and a second data storage unit, and the first data storage unit and the second data storage unit are used alternately; When the mask unit performs the mask operation, it is specifically configured as: If an element of the scale tensor is greater than a preset value, mark the element as a true element, and obtain a corresponding element in the residual tensor based on the element to obtain a retained element pair, and write the retained element pair into the data buffer, where the retained element pair includes a retained residual tensor and a retained scale tensor; Convert the precision of the residual tensor in the retained element pair to a preset precision, and convert the scale tensor in the retained element pair to index information of a preset number of bits; Recombine the retained element pair in the data buffer to output a mask residual tensor, a mask scale tensor, and valid element information.

3. The entropy coding hardware accelerator based on JPEG-AI according to claim 1, characterized in that, The recombined conversion table includes a first recombined conversion table and a second recombined conversion table; among them, the first recombined conversion table is used for residual encoding, and the second recombined conversion table is used for hyper tensor encoding; When the coding rate is changed, the first recombined conversion table remains unchanged, and the second recombined conversion table needs to be switched to the second recombined conversion table corresponding to the coding rate.

4. The entropy encoding hardware accelerator based on JPEG-AI according to claim 3, characterized in that, When the residual encoding module performs encoding, it is specifically configured as: Read the first recombined conversion table and the recombined status table in the on-chip cache unit, and obtain the quantity of the input data, calculate a first value and a first remainder, where the first value is the result of rounding down the input data volume to the nearest multiple of 4, and the first remainder is the remainder generated after performing a division operation on the input data volume by 4; If the first remainder is 0, perform external residual encoding and internal residual encoding on the data of the first value. The external residual encoding is used to perform clipping processing on the elements in the masked residual tensor so that the range of the elements is within a first preset range, and the internal residual encoding is used to perform encoding on the elements after the clipping processing to generate a residual bitstream, and the number of bits of the residual bitstream is the number of bits within a preset range; If the first remainder is non-0, based on the first recombination conversion table and the recombination status table, perform external residual encoding and internal residual encoding in a first preset order to generate a residual bitstream. The first preset order includes a first order and a second order. The first order is the order corresponding to the data of the first remainder quantity, and the second order is the order corresponding to the data of the first value quantity.

5. The entropy encoding hardware accelerator based on JPEG-AI according to claim 3, characterized in that, When the super-tensor encoding module performs encoding, it is specifically configured to: Read the second recombination conversion table in the on-chip cache unit and perform per-channel encoding on the super-tensor based on the second recombination conversion table; Obtain the quantity of the super-tensor and calculate a second value and a second remainder. The second value is the result of rounding down the quantity of the super-tensor to the nearest multiple of 4, and the second remainder is the remainder generated after performing a division operation on the quantity of the super-tensor by 4; If the second remainder is 0, perform external residual encoding and internal residual encoding on the data of the second value. The external residual encoding is used to perform clipping processing on the elements of the super-tensor so that the range of the elements of the super-tensor is within a second preset range, and the internal residual encoding is used to perform encoding on the elements of the super-tensor after the clipping processing to generate a super-tensor bitstream; If the second remainder is non-0, perform external residual encoding and internal residual encoding in a second preset order to generate a super-tensor bitstream. The second preset order includes a third order and a fourth order. The third order is the order corresponding to the data of the second remainder quantity, and the second order is the order corresponding to the data of the second value quantity.

6. The entropy encoding hardware accelerator based on JPEG-AI according to claim 1, characterized in that Both the residual encoding module and the super-tensor encoding module include a storage structure, and the storage structure includes a primary storage structure, a secondary storage structure, and a tertiary storage structure; The width of the primary storage structure is a first width, the secondary storage structure includes a first quantity of first storage units, and the width of the first storage unit is a second width. The tertiary storage structure includes a second quantity of second storage units, and the width of the second storage unit is a third width; The second width is one-half of the first width, and the third width is the product of the first quantity and the second width.

7. The entropy encoding hardware accelerator based on JPEG-AI according to claim 6, characterized in that, The primary storage structure of the residual encoding module and the super-tensor encoding module is used to receive the residual bitstream and the super-tensor bitstream, and shift the code rate in the first direction by the current number of bits, and perform a bitwise OR operation with the data already stored in the slots of the first storage structure to store the residual bitstream; After the second preset quantity of residual encoding and / or super-tensor encoding is completed, obtain the slot status; If the slot status is the half-full status, the secondary storage structures of the residual encoding module and the super-tensor encoding module are used to retrieve and store the data meeting the preset conditions written in the first storage structure, where the preset conditions are the data with high writing priority; Obtain the storage status of the second storage structure; If the storage status is the full status, the tertiary storage structures of the residual encoding module and the super-tensor encoding module read and store the residual bitstream and the super-tensor bitstream stored in the secondary storage structure.

8. The entropy encoding hardware accelerator based on JPEG-AI according to claim 1, wherein Further included are: A bus controller and an off-chip memory. The bus controller communicates with the host through a general protocol and is used to receive the configuration information of the host, where the configuration information includes a start signal, a mode code, address information, and input data scale information; When the bus controller receives the start signal and the mode code is the mask unit mode, the bus controller sends a first start signal to the mask unit to control the mask unit to perform block processing based on the input data and perform a masking operation on the input data; And, based on the address information, the bus controller sends a first control signal to the mask unit, the data loading unit, and the write-back unit to control the data loading unit and the write-back unit to complete the data interaction between the on-chip cache unit and the off-chip memory; When the bus controller receives the first control signal and the mode code is the loading lookup table mode of the data loading unit, the bus controller sends a second control signal to the data loading unit based on the address information to control the data loading unit to load the conversion table and the status table and store them in the on-chip cache unit; And when the bus controller receives the start signal and the mode code is the residual encoding module mode, the bus controller sends a second start signal to the residual encoding module to control the residual encoding module to perform block processing based on the output length of the mask unit and perform an encoding operation; meanwhile, based on the address information, the bus controller sends a second control signal to the residual encoding module, the data loading unit, and the write-back unit to control the data loading unit and the write-back unit to complete the data interaction between the on-chip cache unit and the off-chip memory; The bus controller is further used to send a third start signal to the super-tensor encoding module when the mode code is the super-tensor encoding module mode, control the super-tensor encoding module to perform block processing based on the number of channels and then perform an encoding operation; and based on the address information, the bus controller sends a third control signal to the super-tensor encoding module, the data loading unit, and the write-back unit to control the data loading unit and the write-back unit to complete the data interaction between the on-chip cache unit and the off-chip memory.

9. An entropy coding method based on JPEG-AI, characterized in that, Including: Obtain input data, where the input data includes a residual tensor, a scale tensor, and a super-tensor; Perform a masking operation on the input data to output a masked residual tensor, a masked scale tensor, and valid element information; Recombine the conversion table and the status table to output a recombined conversion table and a recombined status table; Encode the masked residual tensor based on the masked scale tensor, the recombined conversion table, the recombined status table, and the valid element information to output a residual bitstream; And encode the super-tensor based on the recombined conversion table to output a super-tensor bitstream.

Citation Information

Patent Citations

  • Image compression hardware accelerator device based on convolution automatic coding algorithm

    CN111800636A

  • Image coding method, image decompression method and device

    CN115022637A

  • Method and device for encoding image and decoding code stream by using neural network

    CN119547109A

  • Method and system for generating finite state entropy coding table, medium, and device

    WO2023045204A1

  • A method and apparatus for encoding a picture and decoding a bitstream

    WO2025035302A1