JPEG image compression system and method based on FPGA double-matrix sharing assembly line
Through the FPGA dual matrix shared pipeline design, the problem of hardware resource redundancy and coding module isolation in JPEG image compression is solved, efficient YUV three-component data processing is realized, resource occupation and power consumption are reduced, and embedded image compression is suitable for resource-constrained.
Patent Information
- Application Number
- CN202511109247.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-08-08
AI Technical Summary
In the existing JPEG image compression scheme based on FPGA, YUV color space processing has problems such as hardware resource redundancy and coding module isolation, resulting in increased resource usage and waste of resources.
The dual matrix shared pipeline design based on FPGA is adopted, and the shared pipeline is aligned with the timing by a dual 8-row matrix segmentation architecture to realize efficient time-sharing multiplexing of Y, U, and V three-component DCT transformation, Zig-Zag scanning, quantization and Huffman encoding.
It significantly reduces computational complexity and data handling overhead, reduces logic unit and DSP resources, and reduces power consumption and storage requirements, making it particularly suitable for resource-constrained embedded image compression application scenarios.
Smart Images

Figure CN120602650A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing and data compression, and in particular to a JPEG image compression system and method based on an FPGA dual-matrix shared pipeline. Background Art
[0002] Existing image compression algorithms are mainly divided into two categories: lossless compression and lossy compression. Lossless compression does not lose any image information and is achieved simply by encoding the statistical characteristics of the data stream. However, this type of compression generally cannot meet the requirements of high compression ratios. Lossy compression refers to the loss of a certain amount of redundant information or insensitive information during the compression process. It does not have a significant impact on subsequent image processing, but can achieve a high compression ratio. The JPEG compression algorithm is the most widely used lossy compression algorithm. It has high compression efficiency and is simple and easy to implement. It is widely used in digital photography, network transmission, remote sensing communications and other fields.
[0003] Currently, existing FPGA-based JPEG image compression solutions generally use a component-independent processing architecture for YUV color space processing: the luminance component (Y) and chrominance components (U, V) are processed through independent DCT modules, Zig-Zag scanning modules, quantization modules, and Huffman coding modules, respectively. This architecture has the following significant drawbacks: Hardware resource redundancy: Each component needs to be configured with a complete set of processing modules. Taking 8-bit image data as an example, a single DCT module requires 64 multipliers and 192 adders. Independent configuration of the three components will directly multiply the resource usage by 3, seriously increasing the consumption of FPGA logic units (LEs), storage units (RAM) and multipliers (DSP).
[0004] Isolated coding modules: Huffman coding is the core of entropy coding. Existing solutions do not consider the statistical correlation of YUV component code streams. Independent coding leads to duplicate design of code table cache (requiring the storage of three independent code tables) and coding control logic, further wasting resources. Summary of the Invention
[0005] The main purpose of the present invention is to propose a JPEG image compression system and system based on FPGA dual-matrix shared pipeline, which aims to compress image data in YUV color space, realize efficient time-sharing multiplexing of Y, U, and V components for core modules such as discrete cosine transform (DCT), Zig-Zag scanning, quantization and Huffman coding through dual 8-row matrix segmentation architecture and timing alignment shared pipeline design, and solve the technical problem of hardware resource redundancy in traditional solutions.
[0006] To achieve the above-mentioned object, the present invention proposes a JPEG image compression system based on an FPGA dual-matrix shared pipeline, comprising: The YUV420 preprocessing module is used to convert YUV444 format data into YUV420 format data, completely retain the Y component, perform 2:1 horizontal / vertical downsampling on the UV component, and generate a Y component data stream and a 2:1 horizontal / vertical downsampling UV component data stream; A dual 8-row matrix segmentation storage module is used to store the Y component data stream alternately in the Y1 and Y2 buffer areas by row to form a dual 8-row matrix structure, and synchronously store the UV component data stream to obtain Y1, Y2, U, and V data blocks, thereby realizing Y1 / Y2 dual matrix segmentation of the Y component and 8-row aligned storage of the UV component; A shared pipeline timing scheduling module is used to read Y1, Y2, U, and V data blocks from the Y1 and Y2 buffers in the order of Y1 → Y1 → Y2 → Y2 → U → V, generate timing control signals for the shared processing pipeline, implement pipeline-type continuous processing, and convert parallel multi-component data into a serial data stream; A shared processing module is used to perform DCT transformation, quantization, ZigZag scanning, run-length coding and Huffman coding processing of Y1, Y2, U and V data blocks in a time-division multiplexing manner according to the timing control signal; The merge encoding module is used to splice the encoded output into a continuous bit stream and generate output by byte alignment, so as to realize the formatted integration of the encoded data and generate a compressed data stream that conforms to the JPEG standard.
[0007] Optionally, the dual 8-row matrix segmentation storage module generates two 8×8 Y1 and Y2 data blocks alternately every time it receives 16 rows of Y component data through row count control and a segmentation counter. The UV component data stream is cached at an 8-row depth and forms a standard input group of six 8×8 data blocks with the Y1 and Y2 data blocks, providing standardized 8×8 data blocks for the shared pipeline.
[0008] Optionally, the shared pipeline timing scheduling module adopts a three-stage state machine, and the three-stage state machine includes: IDLE state: when the UV component FIFO data volume is detected to be ≥8, it jumps to the read enable generation state; GEN_RD_EN state: Generates the read enable signal in the order of Y1→Y1→Y2→Y2→U→V through the 48-cycle counter; WAIT_BACK state: monitors the backend FIFO empty flag to ensure that there is no data backlog in the shared pipeline.
[0009] Optionally, the shared processing module includes: Shared 2D DCT transform module, used to decompose the 2D DCT into row transform and column transform using the row-column separation algorithm, implement 8-point DCT calculation through a 3-level butterfly network using the Loeffler algorithm, and convert the cosine basis function floating-point coefficients into 16-bit fixed-point numbers; Dynamic quantization module, which is used to perform non-uniform quantization on DCT coefficients based on timing control signals, dynamically switch the luminance / chrominance quantization table, and retain low-frequency information that is sensitive to the human eye; Zigzag pipeline scanning module, used to rearrange the quantized 8×8 quantization coefficient matrix into a one-dimensional sequence according to the ZigZag path, and perform DPCM differential coding on the DC coefficient; The shared run-length encoding module is used to compress the one-dimensional sequence after ZigZag scanning, perform zero-run encoding, and reduce the amount of data; The four tables share the Huffman encoding module, which is used to perform variable-length encoding on the four independent code tables of DC / AC for the Y, U, and V component data after run-length encoding, and assign codes of different lengths according to the probability of data occurrence to achieve lossless compression.
[0010] Optionally, the dynamic quantization module dynamically switches the luminance / chrominance quantization table according to a 48-cycle counter, wherein: Counter in cycles 0-31: Process Y1 and Y2 data blocks and use the brightness quantization table; Counter in cycles 32-47: Processing U, V data blocks and using the chromaticity quantization table.
[0011] On the other hand, the present invention also proposes a JPEG image compression method based on an FPGA dual-matrix shared pipeline, which is performed using the above-mentioned JPEG image compression system based on an FPGA dual-matrix shared pipeline. The JPEG image compression method includes the following steps: Convert YUV444 format data to YUV420 format, generate Y component data stream and UV component data stream after 2:1 horizontal / vertical downsampling; The Y component data stream is alternately stored in the Y1 and Y2 buffer areas by row to form a dual 8-row matrix structure, and the UV component data stream is stored synchronously to achieve Y1 / Y2 dual matrix segmentation of the Y component and 8-row aligned storage of the UV component, providing a standardized 8×8 data block for the shared pipeline; Read data blocks from the Y1 and Y2 buffers in the order of Y1 → Y1 → Y2 → Y2 → U → V, generate timing control signals for the shared processing pipeline, and output them in the order of Y1 → Y1 → Y2 → Y2 → U → V to achieve pipeline continuous processing; Based on the timing control signal, DCT transformation, quantization, ZigZag scanning, run-length coding and Huffman coding processing of Y1, Y2, U and V data blocks are performed in a time-division multiplexing manner; The encoded output is spliced into a continuous bit stream to generate a compressed data stream that conforms to the JPEG standard.
[0012] Optionally, converting the YUV444 format data into the YUV420 format to generate a Y component data stream and a UV component data stream that is downsampled horizontally / vertically by 2:1 comprises the following steps: Data buffer: Build a UV component buffer with a depth of 2 rows, forming a 2×2 pixel window matrix; Average calculation: sum the 4 UV values in the window and perform efficient average calculation by right shifting 2 bits; Data alignment output: Generates a YUV420 format data stream with a Y resolution 2×2 times that of UV, matching the dual 8-row matrix segmentation requirements.
[0013] Optionally, storing the Y component data stream alternately in rows into the Y1 and Y2 buffer areas to form a dual 8-row matrix structure comprises the following steps: Through line count control, every time 16 lines of Y component data are received, two 8×8 Y1 and Y2 data blocks are alternately generated; The UV component data stream is buffered as 8 lines of depth and forms a standard input group of 6 8×8 data blocks with the Y1 / Y2 data blocks.
[0014] Optionally, the time-division multiplexing execution of DCT transformation, quantization, ZigZag scanning, run-length encoding, and Huffman encoding of the Y1, Y2, U, and V data blocks based on the timing control signal comprises the following steps: Based on the 48-cycle counter value, the luminance quantization table is used when processing the Y1 and Y2 data blocks in cycles 0-31, and the chrominance quantization table is used when processing the U and V data blocks in cycles 32-47; Based on four independent code tables for Y / UV components DC / AC, Huffman encoding is performed on different component data.
[0015] Optionally, the time-division multiplexing execution of DCT transforms of the Y1, Y2, U, and V data blocks based on the timing control signal comprises the following steps: Row and column separation: Decompose the two-dimensional DCT into 8 one-dimensional row transforms + 8 one-dimensional column transforms, that is: Among them, f(x,y) is the 8×8 pixel value in the spatial domain, C i (x), C j(y) is a cosine basis function. A one-dimensional DCT unit is called for each row of the 8×8 image block to generate an intermediate frequency domain matrix. The row transform result is converted to column data format by swapping the row and column dimensions through a dual-port BRAM. The one-dimensional DCT unit is then called for the transposed column data to output a complete 8×8 frequency domain coefficient matrix. Butterfly operation: Using the Loeffler algorithm, a three-level butterfly network is used to complete 8-point DCT calculations. Each level requires only four multiplications and eight additions. Utilizing the symmetry of the cosine function, matrix multiplication is decomposed into iterative operations of addition, subtraction, and a small number of multiplications. Fixed-point implementation: Convert the cosine basis function floating-point coefficients to 16-bit fixed-point numbers.
[0016] The technical solution of the present invention has the following beneficial effects: the technical solution of the present invention breaks through the traditional independent module architecture, and realizes efficient time-sharing multiplexing of core modules such as DCT transformation, quantization, ZigZag scanning and Huffman coding of the three YUV components through dual 8-row matrix segmentation and timing alignment shared pipeline design. The computing circuit: compared with the traditional three-channel architecture, it reduces more than 50% of logic units and DSP resources; the storage system: there is no need to cache the intermediate results of the three components at the same time, and the on-chip SRAM requirement is reduced by 30%; power consumption control: the time-sharing multiplexing mechanism reduces the average power consumption by 40% and the data bus bandwidth requirement by 30%; complexity optimization: the data path is simplified and the difficulty of FPGA / ASIC wiring is reduced; the development cycle is shortened by about 25%; the present invention forms a complete and efficient JPEG encoding system by optimizing data organization, processing flow and timing control, which significantly reduces the computing complexity and data handling overhead, and is particularly suitable for resource-constrained embedded image compression application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.
[0018] Figure 1 This is a schematic diagram of the overall module framework structure of a JPEG image compression system based on an FPGA dual-matrix shared pipeline according to an embodiment of the present invention; Figure 2 This is a schematic diagram of a partial module framework structure of a JPEG image compression system based on an FPGA dual-matrix shared pipeline according to an embodiment of the present invention; Figure 3This is a JPEG image compression operation flow chart of a JPEG image compression system based on an FPGA dual-matrix shared pipeline according to an embodiment of the present invention; Figure 4 This is a flow chart of a YUV data storage module of a JPEG image compression system based on an FPGA dual-matrix shared pipeline according to an embodiment of the present invention; Figure 5 The present invention is a flowchart of a shared pipeline timing scheduling module state machine of a JPEG image compression system based on an FPGA dual-matrix shared pipeline according to an embodiment of the present invention.
[0019] Figure 6 A ZigZag path diagram of a JPEG image compression system based on an FPGA dual-matrix shared pipeline according to an embodiment of the present invention; Figure 7 The figure is a schematic diagram of the process steps of a JPEG image compression method based on an FPGA dual-matrix shared pipeline according to an embodiment of the present invention.
[0020] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0022] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative position relationship, movement status, etc. between the various components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.
[0023] In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the fact that ordinary technicians in this field can implement them. When the combination of technical solutions is mutually contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0024] The present invention provides a JPEG image compression system and method based on FPGA dual-matrix shared pipeline.
[0025] like Figures 1 to 6 As shown, in one embodiment of the present invention, the JPEG image compression system based on the FPGA dual-matrix shared pipeline includes: The YUV420 pre-processing module 101 is used to convert the YUV444 format data into the YUV420 format data, completely retain the Y component, perform 2:1 horizontal / vertical downsampling on the UV component, and generate a Y component data stream and a 2:1 horizontal / vertical downsampling UV component data stream; The dual 8-row matrix segmentation storage module 102 is used to store the Y component data stream alternately in the Y1 and Y2 buffer areas by row to form a dual 8-row matrix structure, and simultaneously store the UV component data stream to obtain Y1, Y2, U, and V data blocks, thereby realizing Y1 / Y2 dual matrix segmentation of the Y component and 8-row aligned storage of the UV component; A shared pipeline timing scheduling module 103 is used to read the Y1, Y2, U, and V data blocks from the Y1 and Y2 buffers in the order of Y1 → Y1 → Y2 → Y2 → U → V, generate timing control signals for the shared processing pipeline, implement pipeline-type continuous processing, and convert parallel multi-component data into a serial data stream; The shared processing module 104 is used to perform DCT transformation, quantization, ZigZag scanning, run-length coding and Huffman coding of the Y1, Y2, U and V data blocks in a time-division multiplexing manner according to the timing control signal; The merge encoding module 105 is used to splice the encoded output into a continuous bit stream and generate output according to byte alignment to achieve formatted integration of the encoded data and generate a compressed data stream that complies with the JPEG standard.
[0026] The JPEG image compression system based on FPGA dual matrix shared pipeline of the present invention builds a JPEG compression architecture with dual 8-row matrix segmentation and time alignment shared pipeline as the core, and realizes efficient compression processing of YUV three-component through the collaboration of 9 major modules. The JPEG image compression process is as follows Figure 3 shown.
[0027] Specifically, the YUV420 pre-processing module 101 (dual matrix preparation) is used to implement the format conversion from YUV444 to YUV420, completely retain the Y component, perform 2:1 horizontal / vertical downsampling on the UV component, use 2×2 block mean filtering to ensure chroma smoothing, and provide pre-processed data for dual 8-row matrix segmentation. The specific implementation process steps are as follows: (1) Data cache: Construct a UV component cache with a depth of 2 rows to form a 2×2 pixel window matrix; (2) Mean calculation: Sum the four UV values within the window and perform efficient averaging by right shifting by 2 bits (fixed-point division); (3) Data alignment output: Generates a YUV420 format data stream with a Y resolution of 2×2 times that of UV, matching the dual 8-row matrix segmentation requirements.
[0028] Specifically, the dual 8-row matrix segmentation storage module 102 realizes the Y1 / Y2 dual matrix segmentation of the Y component and the 8-row aligned storage of the UV component through independent caching, row count judgment, write enable triggering and FIFO write control, providing standardized 8×8 data blocks for the shared pipeline, supporting the efficient operation of subsequent encoding, algorithm processing and other links. The realized flow chart is as follows: Figure 4 As shown, the specific implementation process steps are as follows: (1) Parallel channel input: Y, U, and V component data are continuously input through independent channels. When i_de_Y is valid, i_data_Y is written into the Y buffer 8×8 sliding window row by row; similarly, data is written into the U / V buffer 8×8 sliding window through i_de_U / i_data_U and i_de_V / i_data_V. (2) Row count and write enable control: The Y component triggers row count through the rising edge of i_de_Y, and the 3-bit counter row_cnt_Y counts 0-7 in a loop. When row_cnt_Y=7, o_de_Y==i_de_Y, and the counter automatically resets to zero to enter the next row count cycle. The UV component is the same as the Y component. When the row count reaches 7, the output is o_de_U=i_de_U and o_de_V=i_de_V. (3) Y component double buffer division: count the falling edges of o_de_Y and use the 1-bit counter row_cnt_Y (0 / 1 cycle) to implement Y1 / Y2 switching. When row_cnt_Y=0, o_de_Y1=o_de_Y; when row_cnt_Y=1, o_de_Y2=o_de_Y, thereby writing the first 8 rows of Y and the last 8 rows of Y into the corresponding FIFO respectively, realizing the dual 8-row matrix division strategy; (4) Data bit width and transmission cycle: The single pixel bit width is 8 bits / pixel, which complies with the YUV420 standard; the bus bit width is 64 bits (8 pixels × 8 bits / pixel), one line of data is transmitted per clock cycle, and an 8 × 8 data block is written in 8 cycles.
[0029] Specifically, the shared pipeline timing scheduling module 103 is designed to meet the YUV42016×16 matrix operation requirements in JPEG compression processing. It converts parallel multi-component data (Y1 / Y2 / U / V matrices) into a serial data stream and outputs it in the order of Y1→Y1→Y2→Y2→U→V, realizing pipeline continuous processing, eliminating data waiting bottlenecks, and ensuring orderly and efficient processing of each component data.
[0030] The specific implementation process is: use the three-state state machine (IDLE→GEN_RD_EN→WAIT_BACK) to implement the function. The state machine process is as follows: Figure 5 The specific process steps are as follows: (1) Initial state (IDLE) Function: Wait for data to be ready and ensure that there are 8×8 data in both UV component and Y component FIFO.
[0031] Trigger condition: Due to the YUV420 sampling characteristics, the Y component data volume is four times that of the U / V component, and the UV component is written to the FIFO slower than the Y component. Only the amount of data in the U component FIFO needs to be checked. When the U component FIFO data volume is ≥8, the state jumps to the read enable generation state (GEN_RD_EN).
[0032] (2) Read enable generation status (GEN_RD_EN) Function: Generates read enable signals for each component in the order of Y1 → Y1 → Y2 → Y2 → U → V to implement data sequence control, strictly matching the 16×16 matrix processing flow in JPEG compression, ensuring that each 8×8 data block is read completely to avoid data misalignment.
[0033] Trigger condition: A single read requires completing six 8×8 blocks (Y1×2 + Y2×2 + U×1 + V×1), totaling 384 pixels. Using a 64-bit bus with 8 cycles / block, six blocks require 48 cycles, so cnt counts from 0 to 47. When cnt == 47 (48 cycles of reading completed), the trigger state jumps to WAIT_BACK, waiting for the backend FIFO to clear. The specific relationship between the cnt range and read enable control is as follows: (3) Waiting for the backend FIFO to be empty (WAIT_BACK) Function: Serves as a flow control mechanism to prevent backend FIFO overflow, avoid data backlogs caused by slow backend processing speed, and ensure continuous and smooth operation of the pipeline.
[0034] Trigger condition: Monitor the empty flag (empty_back) of the backend FIFO. Only when there is enough space in the backend, return to the initial state and prepare for the next round of reading.
[0035] Specifically, the shared processing module 104 includes a shared two-dimensional DCT transformation module 1041 , a dynamic quantization module 1042 , a ZigZag pipeline scanning module 1043 , a shared run-length coding module 1044 and a four-table shared Huffman coding module 1045 .
[0036] Specifically, the shared two-dimensional DCT transform module 1041 is used to implement a two-dimensional discrete cosine transform (2D-DCT) of an 8×8 image block, converting spatial domain pixels into frequency domain coefficients. Through row-column separation algorithms, the Loeffler algorithm, and fixed-point processing, energy concentration, parallel computing, and resource optimization are achieved. The specific implementation process steps are as follows: (1) Row-column separation algorithm: Decompose the two-dimensional DCT into 8 one-dimensional row transforms + 8 one-dimensional column transforms, namely: Among them, f(x,y) is the 8×8 pixel value in the spatial domain, and Ci(x) and Cj(y) are cosine basis functions.
[0037] A one-dimensional DCT unit is called for each row of the 8×8 image block to generate an intermediate frequency domain matrix. The row transform result is converted to column data format by swapping the row and column dimensions through a dual-port BRAM. The one-dimensional DCT unit is then called again on the transposed column data to output a complete 8×8 frequency domain coefficient matrix. (2) Butterfly operation: Using the Loeffler algorithm, the 8-point DCT calculation is completed through a 3-level butterfly network. Each level only requires 4 multiplications and 8 additions. By utilizing the symmetry of the cosine function, the matrix multiplication is decomposed into an iterative operation of "addition and subtraction + a small amount of multiplication", which reduces the amount of calculation, conforms to the parallel pipeline characteristics of FPGA, and improves processing throughput; (3) Fixed-point implementation: To avoid the high resource consumption of floating-point operations, the floating-point coefficients of the cosine basis function are converted into 16-bit fixed-point numbers. For example, the floating-point coefficient C(0) = 0.3536 is rounded off by 0.3536 × 2^9 = 181.0432 to obtain the fixed-point number 0xB5 (decimal 181).
[0038] Specifically, the dynamic quantization module 1042 (shared table switching) is used to perform non-uniform quantization on the DCT coefficients, dynamically switching the luminance / chrominance quantization table based on a 48-cycle counter to preserve low-frequency information that is sensitive to the human eye. The specific implementation process steps are as follows: (1) Quantization table storage and reading: Each quantization table uses 8×8×8 bit-width data storage, including quantization tables for the luminance component (Y) and chrominance components (U / V). Using a column-first parallel processing architecture, one column of quantization table data (8 elements) is read per clock cycle, and a fixed-point division operation is performed element-by-element with the corresponding column of the DCT coefficient matrix (equivalent to multiplying by the inverse of the quantization table). The quantization operation of the 8×8 matrix is completed in 8 cycles. (2) Quantization table switching mechanism: The system uses a 0-47 loop counter as the timing reference, which is synchronized with the DCT module processing rhythm. Every 8 counts correspond to processing an 8×8 block, and the luminance table (Y) or chrominance table (UV) is dynamically selected based on the counter value: Counter = 0-7: Process the first 8×8 block of the Y1 component, using the luma table.
[0039] Counter = 8-15: Process the second 8×8 block of the Y1 component, using the luma table.
[0040] Counter = 16-23: Process the first 8×8 block of the Y2 component, using the luma table.
[0041] Counter = 24-31: Process the second 8×8 block of the Y2 component, using the luma table.
[0042] Counter = 32-39: Process the first 8×8 block of the U component, using the chrominance table.
[0043] Counter = 40-47: Process the second 8×8 block of the V component, using the chroma table.
[0044] Specifically, the Zigzag pipeline scanning module 1043 is used to implement ZigZag rearrangement of the 8×8 quantization coefficient matrix, DPCM differential encoding, and data pipeline processing, converting the two-dimensional matrix into a one-dimensional sequence, reducing inter-block redundancy, and ensuring continuous data output. The specific implementation process steps are as follows: (1) Matrix cache architecture: Using an 8-level row shift register group, 8 pixels are received and stored in the first row per cycle. Subsequent rows are updated with a delay of 1 cycle through the register chain. After 8 cycles, a complete 8×8 matrix is formed to ensure data timing alignment during ZigZag scanning. (2) ZigZag path decomposition: the matrix is divided into Figure 4 The ZigZag path is split into 8 subsequences, each containing 8 data; (3) Pipeline buffering and output Multi-level cache: Each subsequence is delayed by a register to ensure timing alignment; Timing control: The subsequence is cycled through the counter and the complete sequence is output in 8 cycles; Enable delay: The input enable signal is delayed through 9 levels of registers to ensure synchronization with the data; Data output: processed data is written into FIFO in sequence; (4) FIFO control Read enable trigger: When the amount of remaining data in the FIFO reaches the threshold, the read enable is started; Rhythm control: read enable is controlled by a counter, and one pixel is read per cycle to ensure continuous data output; (5)DPCM processing DC value latch: When the block start position is detected, the current block DC value is latched into the corresponding component register; Differential calculation: calculate the difference between the current DC and the previous DC; Output control: The block start position outputs DC differential, and the remaining positions output AC components; Component switching: Track block types through counters, and process 4 Y blocks, 1 U block, and 1 V block in sequence; Manage each component DC value independently.
[0045] Specifically, the shared run-length encoding module 1044 is used to compress the one-dimensional sequence after ZigZag scanning, representing continuous zero values as "zero run length + non-zero value", reducing the amount of data, especially optimizing the AC component where zero values are concentrated in the high-frequency area. The specific implementation process steps are as follows: (1) Data input: Receive the one-dimensional sequence (DC difference + AC component) after DPCM processing; (2) Zero run counting: When encountering 0, the counter increases by 1; when encountering non-zero, the current count value and non-zero value are output and the counter is cleared; (3) Encoding output: Generate a (run length, non-zero value) combination; output (0,0) at the end of the block as the end of block (EOB); (4) Special processing: All-zero blocks are directly output as EOB; DC differential is output separately and does not participate in zero run counting.
[0046] Specifically, the four-table shared Huffman encoding module 1045 is used to perform variable-length encoding on the run-length-encoded Y, U, and V component data, assigning codes of different lengths based on the probability of data occurrence. Separate Huffman tables are used for the Y component DC / AC and the UV component DC / AC, respectively, to output the code length (size) and binary code (code) corresponding to each symbol, achieving lossless compression. The specific implementation steps are as follows: (1) Loading four sets of Huffman tables: Reading the Y-DC, Y-AC, UV-DC, and UV-AC tables from ROM, which store the mapping relationship between symbols and (code length, encoding); (2) Lookup table encoding: Select the corresponding table according to the component (Y / UV) and type (DC / AC) of the input data; press the combination key to look up the table according to the amplitude value of DC differential and the (run length, non-zero value) of AC to obtain the corresponding code length and binary code; (3) Output control: Output a set of (size, code) per cycle. The code length indicates the number of coded bits, and the binary code is output in a left-aligned format. (4) Special symbol processing: When encountering the end-of-block (EOB), output a fixed code length and code (e.g., size = 2, code = 00); Specifically, the merge coding module 105 is used to splice the code length (size) and binary code (code) output by the Huffman coding module into a continuous bit stream, and generate output according to byte alignment to achieve formatted integration of the encoded data. The specific implementation process steps are as follows: (1) Input buffer architecture design: 32-bit deep buffer registers are used (supporting two splicings of up to 16-bit encoding) with a 5-bit bit counter (recording the current number of buffered bits). During initialization, the buffer is set to 0 and the counter is set to 0 to ensure the continuity of cross-byte encoding; (2) Dynamic bitstream splicing: The code is spliced into the buffer bit by bit according to the code length: the new code is shifted left by the current buffer bit number and merged with the buffer, and the buffer and bit count are updated. It supports up to 16-bit code (compatible with Huffman longest code), and the correct code order is ensured by shifting operations; (3) Byte alignment output control: When the number of bits in the buffer is ≥8, the upper 8 bits are intercepted as byte output, the buffer is updated with the remaining bits, and the counter is decremented by 8. Multi-level register delay is used to ensure that the output data is synchronized with the enable signal to avoid timing conflicts; (4) Cross-byte encoding processing: When the code length exceeds 8 bits, the code is dynamically split through the state machine and conditional judgment: the complete byte is output first, and the remaining bits are retained until the next splicing (for example, if there are 3 bits in the buffer and the new code is 7 bits, the upper 8 bits are output after splicing, and the remaining 2 bits are retained). Predefined case statements are used to handle different combinations of code lengths and remaining bits to ensure that the splicing logic covers all scenarios; (5) Block end and data statistics: When the block end flag is detected, the remaining bits in the output buffer are filled to the byte boundary to generate a block end signal. The total code length data is accumulated and counted for subsequent compression ratio analysis or status monitoring.
[0047] On the other hand, Figure 7 As shown, the present invention also proposes a JPEG image compression method based on an FPGA dual-matrix shared pipeline, which is performed using the above-mentioned JPEG image compression system based on an FPGA dual-matrix shared pipeline. The JPEG image compression method includes the following steps: S100, converting YUV444 format data into YUV420 format, generating a Y component data stream and a UV component data stream that is downsampled horizontally / vertically by 2:1; S200, alternately storing the Y component data stream in Y1 and Y2 buffer areas by row to form a dual 8-row matrix structure, and synchronously storing the UV component data stream, achieving Y1 / Y2 dual matrix segmentation of the Y component and 8-row aligned storage of the UV component, providing a standardized 8×8 data block for the shared pipeline; S300, read data blocks from the Y1 and Y2 buffer areas in the order of Y1 → Y1 → Y2 → Y2 → U → V, generate timing control signals for the shared processing pipeline, and output them in the order of Y1 → Y1 → Y2 → Y2 → U → V to achieve pipeline continuous processing; S400, based on the timing control signal, performing DCT transformation, quantization, ZigZag scanning, run-length coding and Huffman coding processing of the Y1, Y2, U and V data blocks in a time-division multiplexing manner; S500: Splice the coded output into a continuous bit stream to generate a compressed data stream that complies with the JPEG standard.
[0048] Specifically, the conversion of YUV444 format data into YUV420 format to generate a Y component data stream and a UV component data stream that is downsampled horizontally / vertically by 2:1 includes the following steps: Data buffer: Build a UV component buffer with a depth of 2 rows, forming a 2×2 pixel window matrix; Average calculation: Sum the four UV values within the window and perform efficient average calculation by right shifting by 2 bits (fixed-point division); Data alignment output: Generates a YUV420 format data stream with a Y resolution 2×2 times that of UV, matching the dual 8-row matrix segmentation requirements.
[0049] Specifically, the Y component data stream is alternately stored in the Y1 and Y2 buffer areas by row to form a dual 8-row matrix structure, which includes the following steps: Through line count control, every time 16 lines of Y component data are received, two 8×8 Y1 and Y2 data blocks are alternately generated; The UV component data stream is buffered as 8 lines of depth and forms a standard input group of 6 8×8 data blocks with the Y1 / Y2 data blocks.
[0050] Specifically, the DCT transformation, quantization, ZigZag scanning, run-length coding and Huffman coding of the Y1, Y2, U and V data blocks are performed in a time-division multiplexing manner based on the timing control signal, including the following steps: Based on the 48-cycle counter value, the luminance quantization table is used when processing the Y1 and Y2 data blocks in cycles 0-31, and the chrominance quantization table is used when processing the U and V data blocks in cycles 32-47; Based on four independent code tables for Y / UV components DC / AC, Huffman encoding is performed on different component data.
[0051] Specifically, the time-division multiplexing execution of DCT transformation of the Y1, Y2, U, and V data blocks based on the timing control signal includes the following steps: Row and column separation: Decompose the two-dimensional DCT into 8 one-dimensional row transforms + 8 one-dimensional column transforms, that is: Among them, f(x,y) is the 8×8 pixel value in the spatial domain, C i (x), C j (y) is a cosine basis function. A one-dimensional DCT unit is called for each row of the 8×8 image block to generate an intermediate frequency domain matrix. The row transform result is converted to column data format by swapping the row and column dimensions through a dual-port BRAM. The one-dimensional DCT unit is then called for the transposed column data to output a complete 8×8 frequency domain coefficient matrix. Butterfly operation: Using the Loeffler algorithm, a three-level butterfly network is used to complete 8-point DCT calculations. Each level requires only four multiplications and eight additions. Leveraging the symmetry of the cosine function, matrix multiplication is decomposed into iterative operations of addition, subtraction, and a small number of multiplications, reducing the amount of computation, aligning with the parallel pipeline characteristics of the FPGA, and improving processing throughput. Fixed-point implementation: To avoid the high resource consumption of floating-point operations, the floating-point coefficients of the cosine basis function are converted to 16-bit fixed-point numbers.
[0052] Specifically, the basic principles and processes of the technical solution of the present invention are as follows: According to the characteristics of YUV420 data, a vertical matrix segmentation strategy is proposed: (1) Divide the 16-row Y component data into the first 8 rows of Y1 matrix and the last 8 rows of Y2 matrix, and simultaneously extract the 8 rows of U / V components to form independent matrices; (2) Cyclic switching of Y1 / Y2 write enable through a 1-bit row counter generates two 8×8Y matrices, one 8×8U matrix, and one 8×8V matrix every time 16 rows of Y data are received; (3) Beneficial effects: Directly matches 8×8 DCT operations, eliminates traditional row-column conversion overhead, and reduces data cache requirements by 30%.
[0053] 2. Timing alignment shared pipeline Design a dedicated timing control module to achieve orderly scheduling of multi-component data: (1) Use a 48-cycle state machine to output six 8×8 data blocks (Y1×2, Y2×2, U×1, V×1) in the order Y1→Y1→Y2→Y2→U→V; perform DCT transform on each block.
[0054] (2) Trigger data reading based on the UV component FIFO depth threshold to ensure YUV component data synchronization; (3) Shared processing unit: Only one set of DCT transform, quantization, ZigZag scanning, RLE encoding and Huffman encoding modules is required to time-share and multiplex different component data blocks.
[0055] 3. Hardware resource optimization mechanism Efficient resource reuse through architectural innovation: (1) Computing circuit: Compared with the traditional three-channel architecture, it reduces the number of logic units and DSP resources by more than 50%; (2) Storage system: No need to cache the intermediate results of the three components simultaneously, reducing on-chip SRAM requirements by 30%; (3) Power consumption control: The time-sharing multiplexing mechanism reduces average power consumption by 40% and data bus bandwidth requirements by 30%.
[0056] 4. Complete coding process optimization (1) YUV444 to YUV420: 2×2 block mean filtering downsampling, retaining the Y component and compressing the UV component; (2) Two-dimensional DCT transform: row-column separation algorithm combined with Loeffler butterfly network, 16-bit fixed-point implementation; (3) Quantization and Huffman coding: Dynamically switch the brightness / chrominance quantization table, and four independent code tables to implement entropy coding.
[0057] Specifically, the present invention forms a complete and efficient JPEG encoding system by optimizing data organization, processing flow and timing control, which significantly reduces computational complexity and data handling overhead, and is particularly suitable for resource-constrained embedded image compression application scenarios.
[0058] Specifically, compared with the prior art, the present invention has the following advantages: The present invention realizes efficient resource reuse by using the serial processing order of Y1→Y1→Y2→Y2→U→V: (1) Hardware area saving: Only one set of DCT / quantization / ZigZag / Huffman coding processing unit is required (traditional solutions require three sets of parallel units); the computing circuit area is reduced by more than 50%; the control logic is simplified and the timing synchronization complexity is reduced.
[0059] (2) Reduced power consumption: By time-sharing and multiplexing computing units, average power consumption is reduced by about 40%; data bus bandwidth requirements are reduced, reducing storage access power consumption.
[0060] (3) Reduced storage requirements: There is no need to cache the intermediate results of the three components Y, U, and V at the same time; the on-chip SRAM requirement is reduced by about 30%.
[0061] (4) Complexity optimization: data paths are simplified, FPGA / ASIC wiring is reduced, and the development cycle is shortened by approximately 25%.
[0062] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. All equivalent structural transformations made by using the contents of the present invention description and drawings under the inventive concept of the present invention, or direct / indirect application in other related technical fields are included in the patent protection scope of the present invention.
Claims
1. A JPEG image compression system based on FPGA dual-matrix shared pipeline, characterized in that: include: The YUV420 preprocessing module is used to convert YUV444 format data into YUV420 format data, completely retain the Y component, perform 2:1 horizontal / vertical downsampling on the UV component, and generate a Y component data stream and a 2:1 horizontal / vertical downsampling UV component data stream; A dual 8-row matrix segmentation storage module is used to store the Y component data stream alternately in the Y1 and Y2 buffer areas by row to form a dual 8-row matrix structure, and synchronously store the UV component data stream to obtain Y1, Y2, U, and V data blocks, thereby realizing Y1 / Y2 dual matrix segmentation of the Y component and 8-row aligned storage of the UV component; A shared pipeline timing scheduling module is used to read Y1, Y2, U, and V data blocks from the Y1 and Y2 buffers in the order of Y1 → Y1 → Y2 → Y2 → U → V, generate timing control signals for the shared processing pipeline, implement pipeline-type continuous processing, and convert parallel multi-component data into a serial data stream; A shared processing module is used to perform DCT transformation, quantization, ZigZag scanning, run-length coding and Huffman coding processing of Y1, Y2, U and V data blocks in a time-division multiplexing manner according to the timing control signal; as well as The merge encoding module is used to splice the encoded output into a continuous bit stream and generate output by byte alignment, so as to realize the formatted integration of the encoded data and generate a compressed data stream that conforms to the JPEG standard.
2. The JPEG image compression system based on FPGA dual-matrix shared pipeline according to claim 1, characterized in that: The dual 8-row matrix segmentation storage module uses row count control and a segmentation counter to alternately generate two 8×8 Y1 and Y2 data blocks for every 16 rows of Y component data received. The UV component data stream is cached at an 8-row depth and forms a standard input group of six 8×8 data blocks with the Y1 and Y2 data blocks, providing standardized 8×8 data blocks for the shared pipeline.
3. The JPEG image compression system based on FPGA dual-matrix shared pipeline according to claim 1, characterized in that: The shared pipeline timing scheduling module adopts a three-stage state machine, which includes: IDLE state: when the UV component FIFO data volume is detected to be ≥8, it jumps to the read enable generation state; GEN_RD_EN state: Generates the read enable signal in the order of Y1→Y1→Y2→Y2→U→V through the 48-cycle counter; WAIT_BACK state: monitors the backend FIFO empty flag to ensure that there is no data backlog in the shared pipeline.
4. The JPEG image compression system based on FPGA dual-matrix shared pipeline according to claim 1, characterized in that: The shared processing module includes: Shared 2D DCT transform module, used to decompose the 2D DCT into row transform and column transform using the row-column separation algorithm, implement 8-point DCT calculation through a 3-level butterfly network using the Loeffler algorithm, and convert the cosine basis function floating-point coefficients into 16-bit fixed-point numbers; Dynamic quantization module, which is used to perform non-uniform quantization on DCT coefficients based on timing control signals, dynamically switch the luminance / chrominance quantization table, and retain low-frequency information that is sensitive to the human eye; Zigzag pipeline scanning module, used to rearrange the quantized 8×8 quantization coefficient matrix into a one-dimensional sequence according to the ZigZag path, and perform DPCM differential coding on the DC coefficient; A shared run-length encoding module is used to compress the one-dimensional sequence after ZigZag scanning, perform zero run-length encoding, and reduce the amount of data; and The four tables share the Huffman encoding module, which is used to perform variable-length encoding on the four independent code tables of DC / AC for the Y, U, and V component data after run-length encoding, and assign codes of different lengths according to the probability of data occurrence to achieve lossless compression.
5. The JPEG image compression system based on FPGA dual-matrix shared pipeline according to claim 4, characterized in that: The dynamic quantization module dynamically switches the luminance / chrominance quantization table according to a 48-cycle counter, wherein: Counter in cycles 0-31: Process Y1 and Y2 data blocks and use the brightness quantization table; Counter in cycles 32-47: Processing U, V data blocks and using the chromaticity quantization table.
6. A JPEG image compression method based on FPGA dual-matrix shared pipeline, which is performed using the JPEG image compression system based on FPGA dual-matrix shared pipeline according to any one of claims 1 to 5, characterized in that: The JPEG image compression method comprises the following steps: Convert YUV444 format data to YUV420 format, generate Y component data stream and UV component data stream after 2:1 horizontal / vertical downsampling; The Y component data stream is alternately stored in the Y1 and Y2 buffer areas by row to form a dual 8-row matrix structure, and the UV component data stream is stored synchronously to achieve Y1 / Y2 dual matrix segmentation of the Y component and 8-row aligned storage of the UV component, providing a standardized 8×8 data block for the shared pipeline; Read data blocks from the Y1 and Y2 buffers in the order of Y1 → Y1 → Y2 → Y2 → U → V, generate timing control signals for the shared processing pipeline, and output them in the order of Y1 → Y1 → Y2 → Y2 → U → V to achieve pipeline continuous processing; Based on the timing control signal, DCT transformation, quantization, ZigZag scanning, run-length coding and Huffman coding processing of Y1, Y2, U and V data blocks are performed in a time-division multiplexing manner; The encoded output is spliced into a continuous bit stream to generate a compressed data stream that conforms to the JPEG standard.
7. The JPEG image compression method based on FPGA dual-matrix shared pipeline according to claim 6, characterized in that: The conversion of YUV444 format data into YUV420 format to generate a Y component data stream and a UV component data stream that is downsampled horizontally and vertically by 2:1 includes the following steps: Data buffer: Build a UV component buffer with a depth of 2 rows, forming a 2×2 pixel window matrix; Average calculation: sum the 4 UV values in the window and perform efficient average calculation by right shifting 2 bits; Data alignment output: Generates a YUV420 format data stream with a Y resolution 2×2 times that of UV, matching the dual 8-row matrix segmentation requirements.
8. The JPEG image compression method based on FPGA dual-matrix shared pipeline according to claim 6, characterized in that: Storing the Y component data stream alternately in rows into the Y1 and Y2 buffer areas to form a dual 8-row matrix structure includes the following steps: Through line count control, every time 16 lines of Y component data are received, two 8×8 Y1 and Y2 data blocks are alternately generated; The UV component data stream is buffered as 8 lines of depth and forms a standard input group of 6 8×8 data blocks with the Y1 / Y2 data blocks.
9. The JPEG image compression method based on FPGA dual-matrix shared pipeline according to claim 6, characterized in that: The method of performing DCT transformation, quantization, ZigZag scanning, run-length coding and Huffman coding of the Y1, Y2, U and V data blocks in a time-division multiplexing manner based on the timing control signal includes the following steps: Based on the 48-cycle counter value, the luminance quantization table is used when processing the Y1 and Y2 data blocks in cycles 0-31, and the chrominance quantization table is used when processing the U and V data blocks in cycles 32-47; Based on four independent code tables for Y / UV components DC / AC, Huffman encoding is performed on different component data.
10. The JPEG image compression method based on FPGA dual-matrix shared pipeline according to claim 6, characterized in that: The time-division multiplexing execution of DCT transformation of Y1, Y2, U, and V data blocks based on the timing control signal comprises the following steps: Row and column separation: Decompose the two-dimensional DCT into 8 one-dimensional row transforms + 8 one-dimensional column transforms, that is: Among them, f(x,y) is the 8×8 pixel value in the spatial domain, C i (x), C j (y) is a cosine basis function. A one-dimensional DCT unit is called for each row of the 8×8 image block to generate an intermediate frequency domain matrix. The row transform result is converted to column data format by swapping the row and column dimensions through a dual-port BRAM. The one-dimensional DCT unit is then called for the transposed column data to output a complete 8×8 frequency domain coefficient matrix. Butterfly operation: Using the Loeffler algorithm, a three-level butterfly network is used to complete 8-point DCT calculations. Each level requires only four multiplications and eight additions. Utilizing the symmetry of the cosine function, matrix multiplication is decomposed into iterative operations of addition, subtraction, and a small number of multiplications. Fixed-point implementation: Convert the cosine basis function floating-point coefficients to 16-bit fixed-point numbers.
Citation Information
Patent Citations
JPEG (Joint Photographic Experts Group) compression method and device of color digital image
CN101951524A
JPEG compression system based on bin DCT algorithm
CN103491375A
Image compression and decompression method based on FPGA
CN113301344A
Image processing apparatus for compositing images
US20040109610A1