FPGA-based low-delay JPEG-XS multi-stage pipeline discrete wavelet transform processing method, device and system

Through the multi-stage pipeline discrete wavelet transformation method of FPGA, the problems of high latency and large resource consumption of JPEG-XS codec in high-quality video processing are solved, and the image processing effects with low latency, high throughput and low power consumption are achieved.

CN120302053APending Publication Date: 2025-07-11GUILIN UNIV OF ELECTRONIC TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510293962.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing JPEG-XS codecs have problems such as high latency, large resource consumption, insufficient parallel processing capabilities, high hardware design complexity and low resource reuse efficiency when processing high picture quality videos, and it is difficult to meet the performance indicators of the JPEG-XS standard for high throughput, low latency and low power consumption.

Method used

Using a multi-stage pipeline discrete wavelet transformation method based on FPGA, the parallel processing and rapid transmission of image data are achieved through four-pixel parallel input, parity line separation cache, dual-channel SRAM architecture, asymmetric LeGall5/3 transformation module and hybrid boundary processing strategy, and the parallel processing and rapid transmission of image data is achieved, reducing storage requirements and computing time.

Benefits of technology

It realizes low-latency image processing, meets the sub-millisecond delay requirements of the JPEG-XS standard, improves resource utilization and image quality, and maintains high throughput and low power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302053A_ABST
    Figure CN120302053A_ABST
Patent Text Reader

Abstract

The invention discloses an FPGA-based low-delay JPEG-XS multi-stage pipeline discrete wavelet transform processing method, device and system. The method comprises the following steps: carrying out four-pixel parallel input on preprocessed data, separately caching data in odd and even rows, and combining configurable FIFO to obtain a double 5 * 5 pixel processing matrix; based on the asymmetric LeGall5 / 3 conversion module, parallel output of the LL sub-band, the LH sub-band, the HL sub-band and the HH sub-band is achieved after horizontal / vertical double-channel conversion; four-stage horizontal asymmetric LeGall5 / 3 discrete wavelet line transformation is carried out on the output LL sub-band and the output HL sub-band; based on the low-delay transmission requirement of a JPEG-XS intra-frame coding protocol, a four-coefficient parallel output interface based on clock period synchronization is constructed, and a sub-band data structure before entropy coding of JPEG-XS is output. According to the invention, the transmission problem of 8K and 4K ultra-high-definition videos can be solved, the data throughput rate is improved, the transmission delay is reduced, and the method is suitable for scenes with strict requirements on image quality and real-time performance, such as broadcast and TV live broadcast and telemedicine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application belong to the technical field of video encoding hardware acceleration, and particularly relate to a multi-stage pipeline discrete wavelet transform processing method, device and system for low-latency JPEG-XS based on FPGA, which is particularly applicable to the real-time encoding scenario of ultra-high-definition video under the JPEG-XS standard. Background Art

[0002] As a visually lossless encoding standard for ultra-low-latency applications, the core technology of JPEG-XS is based on the LeGall 5 / 3 integer wavelet transform to achieve a compression ratio of 2:1 to 6:1 and a sub-millisecond encoding delay. However, the existing wavelet fast transform methods for JPEG-XS encoding and decoding require wavelet transform on half or a whole image, which results in too high latency of the JPEG-XS codec, consumes a large amount of memory to store intermediate data, and wastes the resources of the FPGA development board. The following technical bottlenecks exist in the FPGA implementation solutions on the market: First, the multi-stage decomposition architecture causes the demand for block memories (BRAMs) to increase exponentially. Processing typical 8K resolution requires more than 1200 BRAM units; second, the single-channel SRAM architecture can only provide an effective bandwidth of 2 pixels / cycle, which cannot meet the parallel processing requirement of 4 pixels / cycle; third, the mirror-symmetric boundary extension algorithm needs to dynamically generate complex address mapping logic, which not only increases the hardware design complexity but also easily introduces periodic artifacts; in addition, the row-column separated filter causes pipeline stage blocking due to data flow dependence, and combined with the repeated calculation units of the multi-stage transform module, the resource reuse efficiency is less than 50%. The above problems jointly restrict the real-time performance and energy efficiency ratio of the ultra-high-definition video processing system, and it is difficult to meet the strict performance indicators of the JPEG-XS standard for high throughput (≥4 pixel / cycle), low latency (≤1 ms frame processing) and low power consumption (<10 W / m 2 ) Summary of the Invention

[0003] In view of the above problems, a multi-stage pipeline discrete wavelet transform processing method, device and system for low-latency JPEG-XS based on FPGA are provided to solve the data throughput bottleneck of the discrete wavelet module in JPEG-XS for 8K and 4K high-quality videos, without the need to process half or a whole image, and reduce the transmission delay by processing the image stream data in a pipeline.

[0004] The present invention provides a method for optimizing the multi-stage pipeline discrete wavelet transform architecture for low-latency JPEG-XS based on FPGA, and the method includes:

[0005] S1. Input the preprocessed data in four-pixel parallel into a dual-channel SRAM architecture for storage, implementing odd-even row separation caching. Combine a four-channel depth-configurable FIFO to generate two 5×5 pixel processing matrices, and output them to an asymmetric LeGall 5 / 3 transform module implemented by a multi-stage pipeline;

[0006] S2. Output the two 5×5 matrices from step S1 to an asymmetric LeGall 5 / 3 transform engine implemented by a multi-stage pipeline. Implement a hybrid boundary processing strategy on the horizontal / vertical dual paths. Among them, the mirror symmetry extension algorithm is applied in the row direction for boundary compensation, and the periodic extension and dynamic phase adjustment mechanism are adopted in the column direction to achieve the parallel output of the LL, LH, HL, and HH subbands. Four coefficients can be processed in parallel in a single clock cycle;

[0007] S3. Based on the two LL and HL subband coefficients output in step S2, alternately input the dual-channel data into a four-stage horizontal asymmetric LeGall 5 / 3 discrete wavelet row transform module through time-division multiplexing to achieve four progressive horizontal decompositions and extract image features with layer-by-layer energy aggregation;

[0008] S4. Based on the low-latency transmission requirements of the JPEG-XS intra-frame coding protocol, construct a four-coefficient parallel output interface synchronized by clock cycles. Through a hardware pipeline architecture, align the bit widths of four discrete wavelet transform coefficients within each clock cycle, and output the subband data structure before the entropy coding of JPEG-XS to achieve a standardized code stream output with a constant coding delay maintained within each pixel sampling period.

[0009] The embodiment of the present application also provides a multi-stage pipeline discrete wavelet transform device for low-latency JPEG-XS based on FPGA. The device includes:

[0010] Input preprocessing module: Used for odd-even row separation caching, supporting four-pixel parallel input. Achieve parallel access to row-level data through a dual-channel design, obtain a dual 5×5 matrix data stream, dynamically adjust the FIFO depth to adapt to different resolutions, and construct a dual 5×5 pixel matrix;

[0011] Asymmetric LeGall 5 / 3 transform module: Implement an asymmetric LeGall 5 / 3 transform module through a multi-stage pipeline, used to perform two-way row and column transforms of the asymmetric LeGall 5 / 3 wavelet on the dual 5×5 matrix data stream to obtain the low-frequency subband LL and high-frequency subbands LH, HL, HH;

[0012] Cache management module: Reuse the same set of transform modules to process different levels of LL subbands, and implement four-level horizontal transform scheduling through a state machine and a counter, used to manage the intermediate data storage of multi-stage wavelet decomposition to obtain the data of multi-stage LL subbands;

[0013] Distribution output module: used to reorganize multi-level sub-band data and compress the output, parallelly output 4 wavelet coefficients per clock cycle, output the sub-band data structure before entropy coding of JPEG-XS, and achieve a standardized bitstream output that maintains a constant coding delay within each pixel sampling period.

[0014] The embodiment of the present application also provides a multi-stage pipelined discrete wavelet transform system for low-latency JPEG-XS based on FPGA, including:

[0015] Programmable on-chip processing unit array;

[0016] An integrated storage module interconnected with the processing unit, used to solidify a computer-executable instruction set;

[0017] When the program is loaded and run on the processing unit, the processor implements the multi-stage pipelined discrete wavelet transform processing method for low-latency JPEG-XS based on FPGA as described above.

[0018] Due to the adoption of the above technical method in the embodiment of the present application, the following technical advantages are achieved: Through the method provided by the present invention, data alignment and boundary extension are performed on the original 4-pixel / period video stream to obtain a double 5×5 matrix data stream; the double 5×5 matrix data stream is used to perform row-column bidirectional transform calculations of the asymmetric LeGall 5 / 3 wavelet to obtain the low-frequency sub-band LL and high-frequency sub-bands LH, HL, HH; the intermediate data is stored through the cache management module to obtain data of multiple levels of LL sub-bands; the multi-level sub-band data is reorganized and compressed for output to obtain data required for JPEG-XS standard entropy coding. Compared with the traditional architecture, the single-frame processing delay of this solution is reduced, meeting the JPEG-XS sub-millisecond requirement and the end-to-end delay being lower than the transmission time of 32 lines; on the premise that the PSNR index remains above 42dB, the image quality is more stable. Description of the Drawings

[0019] Figure 1 It is a schematic flowchart of a multi-stage pipelined discrete wavelet transform processing method for low-latency JPEG-XS based on FPGA according to an embodiment of the present invention;

[0020] Figure 2 It is a structural block diagram of a multi-stage pipelined discrete wavelet transform device for low-latency JPEG-XS based on FPGA according to an embodiment of the present invention; Specific implementation method

[0022] To enable those skilled in the art to better understand the technical solution of the present invention, the embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings of the specification. It should be understood that the embodiments described herein are only presented exemplarily and do not constitute a limitation on the protection scope. Based on the technical solution of the present invention, any technical solution obtained by those skilled in the art through equivalent replacement or reasonable derivation without creative work belongs to the scope of the rights protected by the present invention.

[0023] Referring to Figure 1 , the following is a flowchart of the steps of a multi-stage pipelined discrete wavelet transform processing method for low-latency JPEG-XS based on FPGA according to an embodiment of the present invention, including the following steps:

[0024] S1, The preprocessed data is input in parallel for four pixels into a dual-channel SRAM architecture for storage, realizing odd-even row separation caching. Combining a four-channel depth-configurable FIFO to generate a dual 5×5 pixel processing matrix, and outputting it to an asymmetric LeGall5 / 3 transform module implemented by a multi-stage pipeline.

[0025] This step does not require caching data and accepts the input of the pixel data stream. For the pixel data of the first and last rows and columns of the image, the row-direction mirror symmetry extension technique is used to copy the adjacent two pixel values, and the column-direction periodic extension technique is used to automatically extract the top two and bottom two pixels at the end. A dual 5×5 processing matrix is dynamically generated without pre-storing data, and the video stream data can be processed continuously, which greatly reduces the resources occupied by caching intermediate data.

[0026] S2, The two 5×5 matrices obtained in step S1 are output to an asymmetric LeGall5 / 3 transform engine implemented by a multi-stage pipeline, and a hybrid boundary processing strategy is implemented on the horizontal / vertical dual channels. Among them, the mirror symmetry extension algorithm is applied in the row direction for boundary compensation, and the periodic extension and dynamic phase adjustment mechanism are used in the column direction to realize the parallel output of the LL, LH, HL, and HH subbands, and 4 coefficients can be processed in parallel in a single clock cycle.

[0027] This step implements an asymmetric LeGall5 / 3 transform on a double 5×5 pixel matrix. It should be noted that this data source does not need to store half or one image data, but only needs to calculate data from a double 5×5 pixel matrix. The asymmetric LeGall5 / 3 transform is implemented by a multi-stage pipeline. The data of the asymmetric LeGall5 / 3 transform only requires the data of the double 5×5 pixel matrix transform and the subband data output last time. The row direction uses mirror symmetric extension to eliminate boundary distortion, and the column direction uses periodic extension to construct circular convolution conditions, which greatly reduces the time consumed by discrete wavelet transform; the intermediate coefficient matrix is ​​generated by vertical low-pass / high-pass filtering separation, and then the LL low-frequency subband and HL, LH, and HH three high-frequency subbands are decomposed by secondary filtering in the horizontal direction; the LL subband is recursively input to the lower-level transform unit to realize multi-scale decomposition, and the three high-frequency subbands are compressed and stored in a preset entropy coding format.

[0028] S3, based on the two LL and HL subband coefficients output from step S2, the dual-channel data is alternately input into a four-level horizontal asymmetric LeGall5 / 3 discrete wavelet row transform module through time division multiplexing to achieve four-time progressive horizontal decomposition and realize image feature extraction with layer-by-layer energy aggregation.

[0029] This step reuses the same set of transformation modules to process LL subbands of different levels, and implements four-level horizontal transformation scheduling through a state machine and a counter to manage the intermediate data storage of multi-level wavelet decomposition to obtain data of multi-level LL subbands.

[0030] S4, based on the low-latency transmission requirements of the JPEG-XS intra-frame coding protocol, builds a four-coefficient parallel output interface based on clock cycle synchronization, realizes the bit width alignment of the four discrete wavelet transform coefficients in each clock cycle through the hardware pipeline architecture, and outputs the JPEG-XS entropy coding pre-subband data structure, achieving a standardized bit stream output that maintains a constant coding delay within each pixel sampling cycle.

[0031] This step uses 4 parallel processing units to achieve the synchronous output of 4 wavelet coefficients per clock cycle. This architecture injects the intermediate data stream directly into the next-level pipeline through a phase-interleaved scheduling mechanism, effectively reducing the image frame transmission delay and improving the throughput. It outputs the JPEG-XS entropy coding front subband data structure, achieving a standardized bitstream output that maintains a constant coding delay within each pixel sampling cycle.

[0032] Reference Figure 2 , which is a schematic diagram of the result of a multi-stage pipeline discrete wavelet transform device with low latency JPEG-XS based on FPGA. Figure 2 As shown, the low-latency JPEG-XS multi-stage pipeline discrete wavelet transform device based on FPGA in the embodiment of the present application includes:

[0033] Input preprocessing module 201: It is used for odd-even line separation and caching, supports four-pixel parallel input, realizes parallel access of row-level data through a dual-channel design, obtains a dual 5×5 matrix data stream, dynamically adjusts the FIFO depth to adapt to different resolutions, and constructs a dual 5×5 pixel matrix;

[0034] Asymmetric LeGall 5 / 3 transform module 202: It realizes the asymmetric LeGall 5 / 3 transform module through a multi-stage pipeline, and is used to perform two-way row and column transforms of the asymmetric LeGall 5 / 3 wavelet on the dual 5×5 matrix data stream to obtain the low-frequency sub-band LL and high-frequency sub-bands LH, HL, and HH;

[0035] Cache management module 203: It multiplexes the same set of transform modules to process different levels of LL sub-bands, realizes four-level horizontal transform scheduling through a state machine and a counter, and is used to manage the intermediate data storage of multi-level wavelet decomposition to obtain the data of multi-level LL sub-bands;

[0036] Distribution and output module 204: It is used to recombine multi-level sub-band data and compress and output. Four wavelet coefficients are output in parallel per clock cycle, and the sub-band data structure before entropy coding of JPEG-XS is output to achieve a standardized bitstream output with a constant coding delay maintained within each pixel sampling period.

[0037] Based on the same inventive concept, the embodiment of the present application also provides a multi-stage pipeline discrete wavelet transform system for low-latency JPEG-XS based on FPGA, including:

[0038] A programmable on-chip processing unit array;

[0039] An integrated storage module interconnected with the processing unit, which is used to solidify a computer-executable instruction set;

[0040] When the program is loaded and run on the processing unit, it triggers the processing unit to execute the multi-stage pipeline discrete wavelet transform processing method for low-latency JPEG-XS based on FPGA as described above.

[0041] It should be understood that the exemplary embodiments described herein are illustrative and not restrictive. Although one or more embodiments of the present invention have been described in conjunction with the accompanying drawings, those of ordinary skill in the art should understand that various changes in form and detail can be made without departing from the spirit and scope of the present invention as defined by the appended claims.

Claims

1. A multi-stage pipelined discrete wavelet transform processing method for low-latency JPEG-XS based on FPGA, characterized in that, The steps of the method are as follows: S1. The preprocessed data is input in four-pixel parallel mode to a dual-channel SRAM architecture for storage, achieving odd-even row separation and caching. Combining a four-channel depth-configurable FIFO, a dual 5×5 pixel processing matrix is generated and output to an asymmetric LeGall 5 / 3 transform module implemented by a multi-stage pipeline; S2. Based on the outputs of the two 5×5 matrices in step S1, they are output to an asymmetric LeGall 5 / 3 transform engine implemented by a multi-stage pipeline. A hybrid boundary processing strategy is implemented on the horizontal / vertical dual channels. Among them, the mirror symmetry extension algorithm is applied in the row direction for boundary compensation, and the periodic extension and dynamic phase adjustment mechanism are adopted in the column direction to achieve parallel output of the LL, LH, HL, and HH subbands. Four coefficients can be processed in parallel in a single clock cycle; S3. Based on the two LL and HL subband coefficients output in step S2, the dual-channel data is alternately input into a four-stage horizontal asymmetric LeGall 5 / 3 discrete wavelet row transform module through time-division multiplexing to achieve four progressive horizontal decompositions and realize image feature extraction with energy aggregation layer by layer; S4. Based on the low-latency transmission requirements of the JPEG-XS intra-frame coding protocol, a four-coefficient parallel output interface based on clock cycle synchronization is constructed. Through a hardware pipeline architecture, the bit widths of the four discrete wavelet transform coefficients are aligned within each clock cycle, and the subband data structure before entropy coding of JPEG-XS is output to achieve a standardized bitstream output with a constant coding delay maintained within each pixel sampling period.

2. The multi-stage pipelined discrete wavelet transform processing method of low-latency JPEG-XS based on FPGA according to claim 1, characterized in that, The dual-channel SRAM cache architecture described in step S1 includes an odd-even row separation unit that alternately reads and writes odd-even row data through a dual-port SRAM, and a sliding window generator that dynamically intercepts a dual 5×5 matrix with a 2-pixel step. For the pixel data at the head and tail rows and columns of the image, the mirror symmetry extension technique in the row direction is used to copy the adjacent two pixel values, and the periodic extension technique in the column direction automatically extracts the top two and bottom two pixels at the end to dynamically generate a dual 5×5 processing matrix without prior data storage.

3. The multi-stage pipeline discrete wavelet transform processing method of low-latency JPEG-XS based on FPGA according to claim 1, characterized in that, The data required for S2 comes from the image stream. Without caching half of the image, the discrete wavelet fast transform can be performed, and the end-to-end delay is within 32 rows. S2 performs an asymmetric LeGall 5 / 3 transform on the dual 5×5 matrix. Among them, the mirror symmetry extension is used in the row direction to eliminate boundary distortion, and the periodic extension is applied in the column direction to construct the condition for circular convolution. The required row and column filters are both implemented by multi-stage pipelines, reducing the operation delay of the discrete wavelet; the intermediate coefficient matrix is separated by low-pass / high-pass filtering in the vertical direction, and then the LL low-frequency subband and the three high-frequency subbands of HL, LH, and HH are decomposed by secondary filtering in the horizontal direction; the LL subband is recursively input to the next-level transform unit for multi-scale decomposition, and at the same time, the three high-frequency subbands are compressed and stored in a preset entropy coding format.

4. The multi-stage pipeline discrete wavelet transform processing optimization method for low-latency JPEG-XS based on FPGA according to claim 1, characterized in that, The four horizontal transforms in S3 are implemented through time-division multiplexing in the following way. First, an independent clock cycle is allocated for each level of decomposition, and the computing unit is multiplexed through a state machine. Then, two pipeline registers are inserted between adjacent levels to achieve seamless connection of the data stream.

5. The optimized method for multi - level pipeline discrete wavelet transform processing of low - latency JPEG - XS based on FPGA according to claim 1, characterized in that, S4 eliminates the intermediate data caching bottleneck through a piece pipeline architecture and time-division multiplexing, and uses a 4-way parallel processing unit to synchronously output 4 wavelet coefficients per clock cycle, outputting the sub-band data structure before entropy coding of JPEG-XS, achieving a standardized bitstream output that maintains a constant coding delay within each pixel sampling period.

6. A multi-stage pipelined discrete wavelet transform device for low-latency JPEG-XS based on FPGA, characterized in that, It includes: Input preprocessing module: used for odd-even line separation caching, supporting four-pixel parallel input, realizing parallel access to row-level data through a dual-channel design, obtaining a dual 5×5 matrix data stream, dynamically adjusting the FIFO depth to adapt to different resolutions, and constructing a dual 5×5 pixel matrix; Asymmetric LeGall 5 / 3 transform module: The asymmetric LeGall 5 / 3 transform module is implemented through a multi-stage pipeline, and is used to perform two-way row and column transforms of the asymmetric LeGall 5 / 3 wavelet on the dual 5×5 matrix data stream, obtaining the low-frequency sub-band LL and high-frequency sub-bands LH, HL, and HH; Cache management module: Reusing the same set of transform modules to process different levels of LL sub-bands, realizing four-level horizontal transform scheduling through a state machine and a counter, and being used to manage the storage of intermediate data for multi-level wavelet decomposition, obtaining the data of multi-level LL sub-bands; Distribution and output module: used for reorganizing multi-level sub-band data and compressing and outputting, parallelly outputting 4 wavelet coefficients per clock cycle, outputting the sub-band data structure before entropy coding of JPEG-XS, achieving a standardized bitstream output that maintains a constant coding delay within each pixel sampling period.

7. A low-latency JPEG-XS multi-stage pipelined discrete wavelet transform system based on FPGA, characterized in that, It includes: Programmable on-chip processing unit array; An integrated storage module interconnected with the processing unit, used for solidifying a computer-executable instruction set; When the program is loaded and run on the processing unit, it triggers the processing unit to execute a multi-stage pipeline discrete wavelet transform processing method of a low-latency JPEG-XS based on FPGA according to any one of claims 1 to 5.