ARM (Advanced RISC Machines) and FPGA (Field Programmable Gate Array) collaborative 4K video hard decoding acceleration system and method
Through the 4K video hard decoding acceleration system that cooperates with ARM and FPGA, the efficiency and power consumption problems in traditional technology in 4K video decoding are solved, and the efficient and low-power 4K video real-time decoding is achieved.
Patent Information
- Application Number
- CN202510020809.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-23
AI Technical Summary
When processing 4K resolution video, the decoding process is complex and the hardware resource consumption is high. Traditional CPU and GPU decoding has efficiency and power consumption problems, especially in embedded devices, it is difficult to achieve real-time decoding.
The 4K video hard decoding acceleration system is adopted that coordinates ARM and FPGA. Through the coordinated work of the ARM module and the FPGA module, FPGA is used to efficiently realize core decoding tasks such as CABAC decoding and motion compensation, and combine the AXI bus protocol and Stream FIFO structure to optimize the data handling and caching mechanism.
提高了解码效率,降低解码过程的功耗,支持4K及超高清视频的实时解码,提供高效、低功耗且易于实现的解决方案。
Smart Images

Figure CN120034660A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of digital video technology, and in particular to a 4K video hard decoding acceleration system and method in which ARM and FPGA collaborate. Background Art
[0002] As a new generation of high-efficiency video coding standard, H.265 (HEVC, High Efficiency Video Coding) has been widely recognized in the industry for its remarkable high compression efficiency and excellent video quality. At the same time, thanks to the continuous advancement of ultra-high-definition display technology, 4K resolution video has gradually become the core requirement of mainstream applications. However, the 4K resolution video decoding process is extremely complex, including CABAC entropy decoding, inverse quantization and inverse transformation, intra-frame and inter-frame prediction, motion compensation and other computationally intensive modules, which consumes a lot of hardware resources. Traditional H.265 video decoding is mostly based on general-purpose processors (CPU) or graphics processors (GPU). When processing 4K and above high-resolution videos, it faces multiple limitations. For example, CPU decoding is limited by single-thread performance and has low processing efficiency; GPU has difficulty in optimizing specific modules due to uneven resource utilization. In addition, the high power consumption characteristics of GPU in embedded devices make it difficult to meet the needs of low-power applications, and the high-resolution frame processing delay is large, making it difficult to achieve real-time decoding.
[0003] FPGA technology, with its advantages of high parallel computing, flexible programming capabilities and low power consumption, has gradually become an important technical direction for real-time video processing. However, the current FPGA-based H.265 video decoding solutions are mostly limited to the acceleration optimization of a single module or a few modules, and lack an overall design solution for the full-process decoding architecture, especially in video stream cache management, data transmission, and the collaborative processing mechanism of FPGA and ARM multi-core architecture. There are still many technical difficulties to be overcome. Real-time decoding and transmission of 4K video streams must take into account both high compression rate and low latency requirements. The existing bit rate reduction solutions often lead to reduced encoding efficiency due to the introduction of overly complex algorithms. Summary of the invention
[0004] The purpose of the present invention is to provide a 4K video hard decoding acceleration system and method for ARM and FPGA collaboration, aiming to optimize the 4K video H.265 decoding process through ARM and FPGA collaboration, improve the encoding and decoding efficiency and reduce power consumption.
[0005] To achieve the above object, the present invention provides a 4K video hard decoding acceleration system in which ARM and FPGA collaborate, comprising an ARM module, an FPGA module, an FPGA external DDR and an AXI bus transmission protocol, wherein the ARM module and the FPGA module are connected and communicated via the AXI bus transmission protocol, and the FPGA external DDR is separately connected to the FPGA;
[0006] The ARM module includes a cortex-A9 processor, an AXI communication interface, a DDR storage and an SD card. The cortex-A9 processor is connected to the AXI bus transmission protocol through the AXI communication interface, and the cortex-A9 processor is also bidirectionally connected to the DDR storage and the SD card respectively;
[0007] The FPGA module includes a video decoding module, a video timing control module, an AXI data bit width conversion module, an AXI-Stream FIFO data flow transport IP, and a video format conversion module. The video timing control module is respectively connected to the video decoding module and the AXI bus transmission protocol connection for timing synchronization control. The video decoding module obtains video stream data through the AXI data bit width conversion module and the AXI-Stream FIFO data flow transport IP in turn, and then outputs it through the video format conversion module. One end of the video format conversion module is bidirectionally connected to the FPGA external DDR, and the other end outputs unidirectionally to the AXI bus transmission protocol.
[0008] The video decoding module includes NALU code stream reading IP, NALU parsing IP, NALU Buffer management IP, CABAC decoding IP, inverse quantization and inverse transformation IP, intra-frame prediction IP, inter-frame prediction IP, motion estimation and compensation IP, and filtering IP.
[0009] The NALU code stream reading IP is responsible for parsing the data content of the network abstraction layer unit and passing it to the subsequent decoding module for processing;
[0010] The NALU Buffer management IP is responsible for the network abstraction layer unit data buffer management;
[0011] The NALU parsing IP parses the data output by the NALU Buffer management IP, and according to the provisions of the H.265 standard, parses and outputs the syntax elements corresponding to the code stream;
[0012] The CABAC decoding IP is responsible for decoding the input bitstream, passing the video bitstream to the CABAC decoding module via the data path, and passing the decoded symbols back to the main video decoding process;
[0013] The inverse quantization and inverse transformation IP is responsible for receiving the quantization coefficient and the inverse quantization step as input, performing the inverse quantization operation, and restoring the compressed and encoded video data to the original pixel value;
[0014] The intra prediction IP performs intra prediction on luminance and chrominance, receives the image block and the selected prediction mode as input, and generates prediction values;
[0015] The inter-frame prediction IP performs inter-frame prediction on brightness and chrominance respectively, receives the current frame, reference frame and motion vector as input by means of data interaction with the inverse quantization and inverse transformation IP and the motion estimation and compensation IP, generates a predicted frame, adds the predicted frame and the residual decoded by the inverse transformation and inverse quantization module, obtains a decoded frame, and finally sends the decoded frame to the filtering module for filtering;
[0016] The motion estimation and compensation IP is used to calculate motion vectors and perform motion compensation;
[0017] The filtering IP includes a SAO filtering module and a deblocking filtering module. The deblocking filtering module receives the decoded image data, identifies boundary pixels, and applies a filtering algorithm to reduce the blocking effect; the SAO filtering module receives the decoded image data and adjusts the pixel value according to the neighborhood information around the pixel to reduce artifacts.
[0018] The AXI bus transmission protocol includes three different interface types: AXI-Lite, AXI4 and AXI-Stream.
[0019] The DDR storage includes a DDR controller and a DDR SDRAM memory chip, which is used for reading cache of the video stream to be decoded and local cache of the video stream after decoding.
[0020] The SD card is used to store the video stream to be decoded and the video stream after decoding, and the reading and writing of data are controlled by the cortex-A9 processor.
[0021] The present invention also proposes a 4K video hard decoding acceleration method in which ARM and FPGA are coordinated. Using the 4K video hard decoding acceleration system in which ARM and FPGA are coordinated, the method comprises the following steps:
[0022] Step 1: The system is powered on, and the cortex-A9 processor in the ARM module initializes the hardware parameter configuration of the FPGA module;
[0023] Step 2: The ARM module loads the H.265 video data stored on the SD card to the DDR storage. The FPGA module accesses the DDR SDRAM memory chip of the ARM module through the AXI-Stream FIFO data flow transport IP, reads the H.265 image data, and writes it to the FPGA external DDR after data format and bit width conversion. The video stream decoding module starts working;
[0024] Step 3: The video decoding module reads the FPGA external DDR to obtain the video stream to be decoded, reads the NALU stream data through the NALU stream reading IP, manages the buffer of the stream through the NALU parsing IP and the NALU Buffer management IP, and parses the stream, respectively parsing the three parsing IPs of VPS, PPS, and SPS, and then uses the CABAC decoding IP to decode the three bins, and divides the decoded data into 3 paths, one path is sent to the dequantization IP and the inverse transformation IP for inverse transformation and quantization, the other path is sent to the intra-frame prediction IP, and the last path is sent to the inter-frame prediction IP and the motion compensation IP, and finally the output video image is obtained through the filtering IP;
[0025] Step 4: After obtaining the decoded video image, the decoded video image is sent to the DDR SDRAM memory chip of the ARM module for storage. After the processor reads the DDR end data, it writes the decoded YUV420 video image into the SD card for storage.
[0026] Optionally, the AXI-Lite of the AXI-Stream FIFO data flow handling IP in step 2 is used for ARM to configure the video stream data transmission channel of AXI-Stream, including the length of the data to be transmitted and the occupancy of the FIFO;
[0027] The AXI_STR_TXD interface of the AXI-Stream FIFO data stream handling IP is used to handle the data transmitted from the ARM DDR and convert it into video stream data in AXI-Stream format.
[0028] Optionally, VPS is a video parameter set, and the VPS parsing IP module is used to extract parameter information describing the decoder buffer and video sequence characteristics in the video stream;
[0029] PPS stands for picture parameter set, and the PPS parsing IP module focuses on parsing the picture parameter set, especially for fine-grained control of the slice layer;
[0030] SPS is a sequence parameter set. The SPS parsing IP module further parses the sequence parameter set in the NALU code stream, including the video frame width, height, chroma format, and maximum encoding block size parameters.
[0031] The present invention provides a 4K video hard decoding acceleration system and method for ARM and FPGA collaboration. The system is composed of an ARM module, an FPGA module, an FPGA external DDR, and an AXI bus transmission protocol. Specifically, the ARM module and the FPGA module are used to coordinate advantages to comprehensively optimize the calculation and transmission process in the 4K video H.265 decoding process, and the FPGA module is used to efficiently implement core decoding tasks such as CABAC decoding and motion compensation modules. At the same time, the AXI bus protocol and the Stream FIFO structure are used to optimize data handling and caching mechanisms to improve decoding efficiency. Compared with the pure software decoding solution, the present invention improves the decoding efficiency, combines the collaborative optimization of FPGA and ARM processors, and effectively reduces the power consumption during the decoding process. In addition, through modular design and flexible configuration of FPGA resources, the present invention can support expansion to other video standards (such as H.264) or higher resolution decoding requirements, and provides an efficient, low-power and easy-to-implement solution for real-time decoding of 4K and ultra-high-definition video. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0033] Figure 1 It is a schematic diagram of the structural principle of a 4K video hard decoding acceleration system in which ARM and FPGA collaborate.
[0034] Figure 2 It is a schematic diagram of a partial FPGA engineering architecture of an H.265 video hard decoding system according to a specific embodiment of the present invention.
[0035] Figure 3 It is a structural diagram of an AXI-Stream FIFO data stream transport IP module according to a specific embodiment of the present invention.
[0036] Figure 4 It is a structural diagram of a NALU code stream reading IP module according to a specific embodiment of the present invention.
[0037] Figure 5 It is a structural diagram of a NALU Buffer management IP module according to a specific embodiment of the present invention.
[0038] Figure 6 It is a structural diagram of a NALU parsing IP module according to a specific embodiment of the present invention.
[0039] Figure 7It is a schematic diagram of the structure of a CABAC decoding IP module according to a specific embodiment of the present invention.
[0040] Figure 8 It is a structural diagram of the inverse quantization and inverse transformation IP module of a specific embodiment of the present invention.
[0041] Fig. 9 It is a structural diagram of an intra-frame prediction IP module according to a specific embodiment of the present invention.
[0042] Fig.10 It is a structural diagram of the inter-frame prediction IP module according to a specific embodiment of the present invention.
[0043] Fig.11 It is a structural diagram of a motion estimation and compensation IP module according to a specific embodiment of the present invention.
[0044] Fig.12 It is a schematic diagram of the structure of the filtering IP module according to a specific embodiment of the present invention.
[0045] Fig.13 The present invention is a control flow diagram of a 4K video hard decoding acceleration method in which ARM and FPGA collaborate. DETAILED DESCRIPTION
[0046] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.
[0047] See also Figure 1 , the present invention provides a 4K video hard decoding acceleration system in which ARM and FPGA cooperate, including an ARM module, an FPGA module, an FPGA external DDR and an AXI bus transmission protocol, wherein the ARM module and the FPGA module are connected and communicated via the AXI bus transmission protocol, and the FPGA external DDR is separately connected to the FPGA;
[0048] The ARM module includes a cortex-A9 processor, an AXI communication interface, a DDR storage and an SD card. The cortex-A9 processor is connected to the AXI bus transmission protocol through the AXI communication interface, and the cortex-A9 processor is also bidirectionally connected to the DDR storage and the SD card respectively;
[0049] The FPGA module includes a video decoding module, a video timing control module, an AXI data bit width conversion module, an AXI-Stream FIFO data flow transport IP, and a video format conversion module. The video timing control module is respectively connected to the video decoding module and the AXI bus transmission protocol connection for timing synchronization control. The video decoding module obtains video stream data through the AXI data bit width conversion module and the AXI-Stream FIFO data flow transport IP in turn, and then outputs it through the video format conversion module. One end of the video format conversion module is bidirectionally connected to the FPGA external DDR, and the other end outputs unidirectionally to the AXI bus transmission protocol.
[0050] In this embodiment, the DDR3 / DDR4 SDRAM memory chip of the ARM module processor is used for reading cache of the video stream to be decoded and local cache of the video stream after decoding. The chip is connected to an SD card for storing the video stream to be decoded and the video stream after decoding, and the reading and writing of data are controlled by the ARM module processor. The FPGA module is used to carry the video stream data, then decode the video stream, and write the decoded video stream data back to the SD card.
[0051] The following is further described in conjunction with the structure of a specific embodiment (H.265 decoder):
[0052] See also Figures 2 to 12 , the ARM-FPGA collaborative H.265 decoder is verified on the ZYNQ7035 platform. The ARM module is responsible for loading the video data to be decoded and storing the decoded video data; the FPGA module is responsible for decoding the video data, mainly including the stream interface for transmitting the bitstream to the decoder, the AXI high-speed transmission interface HP for writing the decoded YUV data, and reading the reference frame. Figure 2 As shown, it includes AXI-Stream FIFO data stream handling IP, video timing control, H.265 decoding and other IPs.
[0053] like Figure 3As shown, it is a schematic diagram of the data stream handling and format conversion IP. The data stream handling and format conversion IP directly calls the AXI-Stream FIFO IP of XILINX. The AXI4-Stream FIFO IP allows memory-mapped access to the AXI4-Stream interface. This IP can be used to connect to the AXI4-Stream IP, connect the H.265 decoder with the help of the AXI_STR_TXD interface, and transmit the AXI-Stream video stream to be decoded to the decoder without using a complete DMA solution. The S_AXI interface is a memory-mapped interface for accessing configuration registers and data Tx and Rx data, which can be used for ARM to configure the video stream data transmission channel of AXI-Stream, including the length of the data to be transmitted, the occupancy of the FIFO, etc. The main operation of this IP allows data packets to be written to or read from the device without considering the AXI4-Stream interface signaling, and the axis4-stream interface can be easily managed because they are transparent.
[0054] like Figure 4 The figure shows a schematic diagram of the NALU code stream reading IP. The NALU code stream reading IP is mainly responsible for reading the video byte stream to be decoded, outputting the divided code stream type, and passing the data to the next module for processing.
[0055] Similar to the H264 / AVC structure, H265 / HEVC also uses the video coding layer (VCL) and the network abstract layer (NAL). The VCL layer contains the video compression data, and the NAL is mainly responsible for dividing and encapsulating the compressed data to ensure that the data is stored on the disk and transmitted on the network.
[0056] The NALU code stream reading IP module proposed in the present invention is designed for H.265 / HEVC video decoding, and is intended to efficiently and accurately parse the data content of the network abstraction layer unit (NALU) and pass it to the subsequent decoding module for processing. The module receives the input H.265 code stream data through the standard AXI interface, uses the start code detection mechanism to identify the 0x000001 or 0x00000001 sequence in the code stream, divides it into independent NAL units, and further parses the type information of each NALU (such as video frames, parameter sets, etc.). During the parsing process, the module reads the bits of the type field after the start code according to the NALU format of the H.265 standard, converts it into a specific NALU type code, and outputs the parsed data to the subsequent module through the internal buffer. In addition, the NALU code stream reading IP has a robust error detection and recovery mechanism, which can effectively deal with the problem of start code loss or code stream damage that may occur during the transmission process, and ensure the reliability of the parsing results. In terms of structural design, the module integrates multiple input and output interfaces, including clock and reset signals (clk and rst), module enable signal (en), data read request signal (rd_req_by_rbsp_buffer_in), and input code stream data interface (mem_data_in[7:0]). The NALU type parsed by the module is transmitted through the output interface nal_unit_type[5:0], and the parsed data stream is output through the interface rbsp_data_out[7:0], with the validity flag signal rbsp_valid_out. At the same time, the module indicates the location of the next NAL unit through the next_nalu_detected signal, realizing the smoothness and real-time performance of data parsing.
[0057] like Figure 5 The figure shows a schematic diagram of the NALU Buffer management IP. NALU buffer management is a core component of the decoder, ensuring that video data is correctly received, stored and processed. The NALU Buffer management IP module proposed in the present invention is an important hardware module for implementing network abstraction layer unit (NALU) data buffer management in the H.265 / HEVC decoder, and is intended to provide efficient and reliable basic support for the subsequent parsing of video parameter sets (VPS), sequence parameter sets (SPS) and picture parameter sets (PPS). NALU buffer management is an indispensable core link in the H.265 video decoding process, involving the dynamic storage and management of NALU data units, ensuring that the decoder can accurately and orderly process the data stream input from the NALU code stream reading IP module.
[0058] This module is implemented in hardware and is divided into two parts: rbsp_buffer_simple_0 module and rbsp_buffer module. The rbsp_buffer_simple_0 module is responsible for the initial buffering of NALU code stream data, receiving the RBSP data stream of the NALU code stream through the input interface, and combining the forward length (forward length) and NALU detection signal for caching, data format correction and validity verification to ensure the integrity and continuity of data output; its output interface sends the corrected NALU code stream to the rbsp_buffer module for further management. The rbsp_buffer module focuses on the efficient management of the NALU buffer, supports dynamic adjustment of the buffer size to adapt to the storage requirements of different data units, and realizes the orderly output of data through zero bit check, last bit flag and read request mechanism. Flexible zero bit counting and NALU type parsing functions are introduced in the design of this module to ensure that the output data can be accurately connected to the subsequent parsing module to meet the parsing requirements of various NALU types.
[0059] like Figure 6 The figure shows a schematic diagram of the NALU parsing IP. The NALU parsing IP is divided into three parsing IPs: VPS (video parameter set), PPS (picture parameter set), and SPS (sequence parameter set). The data output by the NALU Buffer management IP is parsed. According to the provisions of the H.265 standard, the syntax elements corresponding to the code stream are parsed and output to other modules.
[0060] For a bitstream file, like H264, there are a series of NALU type definitions, which can be divided into 6 types: VPS, SPS, PPS, SEI, I frame, and P frame. The bitstream structure is as follows:
[0061] Start code + VPS + start code + SPS + start code + PPS + start code + SEI + start code + I frame + start code + P frame + start code + P frame + ..... The decoder can decode H265 data only after getting VPS, SPS, PPS. The H.265 code stream always starts with VPS, SPS, PPS.
[0062] The VPS parsing IP module is used to extract parameter information describing the decoder buffer and video sequence characteristics in the video stream. Its interface includes input RBSP code stream, clock and reset signals, as well as output maximum decoded picture buffer number, status signal and data forward length, etc., which provides basic support for the decoder to correctly allocate resources and manage cache. Through efficient hardware implementation, the VPS module can parse complex VPS fields, such as vps_max_dec_pic_buffering, to ensure accurate calculation and delivery of cache requirements in video streams.
[0063] The SPS parsing IP module further parses the sequence parameter set in the NALU code stream, including the video frame width, height, chroma format, maximum coding block size and other parameters. These parameters directly affect the decoder's image reconstruction and block processing strategy. The hardware design of the SPS module supports fast parsing of multiple key fields, such as log2_max_pic_order_cnt_lsb_minus4 for managing the display order, log2_min_coding_block_size_minus3 and log2_diff_max_min_coding_block_size for defining the range of the minimum and maximum coding blocks. In addition, the SPS module also outputs a series of flags related to prediction modes and filters, such as sample_adaptive_offset_enabled_flag and strong_intra_smoothing_enabled_flag, which provide guarantees for improving decoding quality and efficiency.
[0064] The PPS parsing IP module focuses on the parsing of image parameter sets, especially the fine-grained control of slices, such as quantization parameter offset, decoder parallelism setting, and deblocking filter configuration. Its interface includes input support for NALU bitstreams and output of multiple parsing results, such as pps_cb_qp_offset and pps_cr_qp_offset for color channel quantization adjustment, deblocking_filter_control_present_flag and loop_filter_across_slices_enabled_flag for cross-slice filtering control. The efficient parsing capability of the PPS parsing module can provide the decoder with accurate image-level parameters, enabling high-quality video reconstruction in complex encoding scenarios.
[0065] like Figure 7 The figure shows a schematic diagram of the CABAC decoding IP. The CABAC decoding IP includes functions such as context modeling, binary bit decoding, syntax element parsing, and context updating. It is responsible for decoding the input bit stream, passing the video bit stream to the CABAC decoding module with the help of the data path, and passing the decoded symbols back to the main video decoding process. At the same time, a state machine is created for the CABAC decoder to control the various stages of CABAC decoding, including the update of the context model and the decoding of symbols. In order to further optimize the FPGA design to improve performance, including through pipelining, parallel processing, etc., as well as to reduce critical paths and reduce power consumption, it is divided into 5 blocks. After decoding, the data is passed to the inverse transform and inverse quantization IP.
[0066] Context-Based Adaptive Binary Arithmetic Coding (CABAC) is a context-based algorithm for encoding symbols in a video into a binary bit string. CABAC decoding is the first step in the H.265 video decoding framework, which is used to restore the binary data used in video encoding. This is a complex process that usually includes the following steps:
[0067] ① Context modeling: CABAC uses different context models to represent different symbols. The context model is a model that determines how to encode the current symbol based on the encoded data context and a set of state information.
[0068] ② Binary bit decoding: CABAC decodes each symbol into a binary bit string. Each bit is determined by the CABAC decoder based on the current context model.
[0069] ③ Syntax element parsing: The decoder also needs to understand the syntax elements represented by the decoded bit string, such as motion vectors, quantization parameters, etc.
[0070] ④Reconstruction: The decoded symbols and syntax elements are used to reconstruct the video frames.
[0071] ⑤Context update: After decoding is completed, the context model needs to be updated so that it can better predict when decoding the next symbol.
[0072] Its interface design covers a variety of input signals, including bitstream input i_rbsp_in[7:0], slice quantization parameter i_sliceQpY[5:0], slice type i_slice_type[1:0] and related CABAC initialization flags, such as i_cabac_init_present_flag and i_cabac_init_flag, which are used to control the context initialization process of the decoder. In addition, the module receives index values of the context model, such as i_cm_idx_cu[4:0], i_cm_idx_sd[2:0], etc., to support context modeling of different syntax elements. The core logic of CABAC decoding relies on the dynamic adjustment of the current range value o_ivCurRange_r[8:0] and the offset value o_ivOffset_r[8:0] to complete the arithmetic decoding operation.
[0073] The output interface includes multiple binary results for decoding syntax elements, such as o_bin_cu (for CU-related syntax), o_bin_sd (for SD-related syntax), o_bin_sig (significance flag), etc. At the same time, the o_bin_byp and o_bin_term interfaces support the decoding of bypass mode and termination bits to meet the processing requirements of special syntax elements. The module also provides control signals o_valid and output data length o_output_len[2:0] to ensure the correctness and timing consistency of the decoded data.
[0074] like Figure 8 The figure shows the schematic diagram of the inverse quantization and inverse transform IP. Determine the required input and output interfaces, as well as the internal data path. Implement the inverse quantization module, which receives the quantization coefficient and the inverse quantization step size as input and performs the inverse quantization operation. And design the hardware multipliers and adders to perform multiplication and addition operations with the help of multipliers and adders. Depending on the type of transform used (such as DCT or DST), implement the inverse transform module, including the calculation of the inverse transform matrix and the inverse transform operation.
[0075] In the video decoding process, inverse quantization and inverse transform are one of the key steps in the decoding process. The purpose of these two steps is to restore the compressed and encoded video data to the original pixel values. During the encoding process, H.265 uses a quantizer to reduce the precision of the transform coefficients and represent them with fewer bits. In the decoding process, inverse quantization is required to restore these coefficients. The inverse quantization process involves multiplying the quantized coefficients by the inverse quantization step size, which is usually determined by the quantization parameter in the encoding process.
[0076] In H.265, multiple transforms are commonly used, such as discrete cosine transform (DCT) and discrete sine transform (DST). During the decoding process, an inverse transform is required to restore the transform coefficients to spatial pixel values. The exact type of inverse transform depends on the transform type selected during the encoding process. Typically, the inverse transform process involves multiplying the transform coefficients by the inverse transform matrix and then rounding to integer pixel values.
[0077] The Dequantization and Inverse Transformation IP modules are used to restore quantization coefficients and transform domain data in the H.265 / HEVC video decoding process. These modules are responsible for decoding the quantized transform coefficients in the compressed video stream into pixel values in the spatial domain, which is one of the key steps in reconstructing the video. The Dequantization and Inverse Transformation IP is divided into two modules: trans_quant_16 and trans_quant_32, which perform precise dequantization and inverse transformation operations on transform units (TUs) of different sizes.
[0078] The trans_quant_16 module mainly processes transform units of size 16×16 and below. The interface design includes clock and reset signals (such as clk, rst, global_rst), control signals (such as en, i_transfom_skip_flag, i_transquant_bypass), and input and output data path signals. The module processes the quantization coefficients block by block by receiving the transform unit size parameters (such as i_log2TrafoSize, i_trafoSize), quantization parameters i_qp[5:0], prediction mode i_predmode, etc., and generates inverse quantized pixel values. Input signals such as bram_coeff_addr and bram_coeff_dout are used to read the quantization coefficients from the memory, while output signals such as dram_tq_we, dram_tq_addrd and dram_tq_did write the processed results to the next processing module. In addition, the module also includes a state signal o_trans_quant_state and a completion flag o_tq_done_y, which are used to indicate the processing progress of the inverse quantization and inverse transformation.
[0079] The trans_quant_32 module is similar in design, but is extended for larger transform units (32×32 and above). Its interface design further optimizes data throughput and hardware resource utilization while supporting a larger input range. The module receives input parameters such as the transform unit starting coordinates i_x0, i_y0, transform size and quantization parameters, and implements block-level processing of large-size data by configuring signals i_xTu, i_yTu, i_x_Imt and i_y_Imt. The dequantized coefficients are output to subsequent modules through the data path interface, and can also be written back to the memory for further processing. The module design also considers special processing of bypass mode and skip transform mode, so that it can adapt to a variety of encoding scenarios, further improving the robustness of the decoder.
[0080] like Fig. 9 As shown in the figure, it is a schematic diagram of the intra prediction IP. The intra prediction IP performs intra prediction on luminance (luma) and chrominance (chroma), receives the image block and the selected prediction mode as input, and generates a prediction value. A submodule is designed for each possible prediction mode, and one of them is selected according to the selected mode. For each block, the residual between the actual pixel value and the predicted value is calculated. The calculated residual is encoded and added to the bitstream. This usually requires the use of entropy coding techniques defined in video coding standards, such as Huffman coding or context-adaptive binary arithmetic coding (CABAC).
[0081] The H.265 / HEVC video coding standard uses intra-frame prediction to exploit the spatial redundancy of images within a frame to further reduce the size of video data. The intra-frame prediction process allows the encoder to predict pixel values within the image and store only the residual, thereby achieving higher compression rates.
[0082] First, the video frame is divided into multiple small blocks, usually square, such as 4x4, 8x8 or 16x16 pixel blocks. For each block, the intra prediction process will make a prediction based on its neighboring pixel values.
[0083] For each block, an appropriate prediction mode is selected. H.265 defines 35 prediction modes, including horizontal, vertical, DC (mean), angle, and various combination modes. The goal of selecting the best mode is to minimize the prediction residual.
[0084] The selected forecasting mode is used to generate forecast values. The specific forecasting process depends on the selected mode.
[0085] The residual (difference) between the actual pixel value and the intra-frame predicted value is calculated. The residual represents the error of the selected prediction mode, which usually includes positive and negative values. The encoder encodes the residual and adds it to the bitstream.
[0086] The intra prediction IP is divided into two sub-modules, namely the intra_pred_16 module for small-sized transform units and the intra_pred_32 module for large-sized transform units, which together complete the prediction processing for different block sizes.
[0087] The intra_pred_16 module provides flexible and efficient intra prediction functions for transform units of 16×16 and below. Through the input interface, the module receives the position information of the transform unit (such as i_x0 and i_y0), block size parameters (such as i_log2TrafoSize and i_trafoSize), neighborhood pixel buffers (such as i_line_buf_left, i_line_buf_top and i_leftup), and intra prediction mode i_intra_predmode. According to the H.265 / HEVC standard, the module supports multiple prediction modes, including planar prediction, DC prediction and up to 35 directional predictions, and can generate predicted pixels based on the available pixel values around the block. The module output interfaces such as dram_pred_we, dram_pred_addr and dram_pred_din are used to write the prediction results to the next stage storage, and the prediction completion status and internal state transition are indicated through o_pred_done_y and o_intra_pred_state.
[0088] The intra_pred_32 module extends the processing capacity of intra-frame prediction and is applicable to transform units of 32×32 and above. Similar to the intra_pred_16 module, its design includes rich interfaces to receive and process a wider range of input data, such as larger neighborhood caches i_line_buf_left and i_line_buf_top. The module supports Strong Intra Smoothing, and the signal i_strong_intra_smoothing_enabled_flag controls whether to enable this function. For the prediction of larger-sized blocks, the module provides higher data throughput capacity, supports a larger range of storage operations (such as higher-bitwidth dram_pred_din and dram_pred_addr), and adopts a pipeline architecture to ensure prediction efficiency.
[0089] As Fig.10 shown, it is a schematic diagram of the inter-frame prediction IP. The inter-frame prediction IP performs inter-frame prediction on luminance (luma) and chrominance (chroma) respectively. By means of data interaction with the inverse transform and inverse quantization IP and the motion estimation and compensation IP, it receives the current frame, reference frame, and motion vector as inputs, generates a prediction frame. The prediction frame is added to the residual decoded by the inverse transform and inverse quantization module to obtain the decoded frame. Finally, the decoded frame is sent to the filtering module for filtering.
[0090] The H.265 / HEVC video coding standard uses inter-frame prediction to utilize the temporal redundancy between adjacent frames in a video sequence to further reduce the size of video data. The inter-frame prediction process allows the encoder to predict the pixel values of the current frame by referring to previous frames (reference frames), thereby reducing storage and transmission overhead.
[0091] By interacting with the motion estimation and compensation module, a motion vector is obtained. The motion vector is used to perform motion compensation, moving the pixels in the reference frame to the position of the current frame to generate a prediction frame. The prediction frame is added to the residual decoded by the inverse transform and inverse quantization module to obtain the decoded frame.
[0092] Inter-frame prediction is a key component of H.265 / HEVC video decoding. It is used to utilize the temporal correlation between the current frame and the reference frame, predicting the pixel values of blocks through motion compensation technology to reduce data redundancy. The inter-frame prediction IP is mainly divided into two modules, which perform prediction processing on luminance (luma) and chrominance (chroma) respectively to meet the requirements of different components in terms of prediction accuracy and data transmission.
[0093] The inter_pred_chroma module is mainly used to process inter-frame prediction of chroma components (such as Cb and Cr). By receiving input signals such as block position i_x0 and i_y0, block size i_nPbW and i_nPbH, motion vectors i_mvx and i_mvy of the reference frame, the module can accurately locate the prediction area of the chroma component. The module supports dynamic adjustment of the start and end positions of the reference area, and specifies the prediction area range through the output signals o_ref_start_x, o_ref_end_x, o_ref_start_y and o_ref_end_y. In addition, the module designs an efficient FIFO data management mechanism, combined with the i_fifo_data and i_fifo_empty interfaces to realize the orderly storage and reading of prediction data, ensuring the integrity and timing synchronization of the data. After the prediction is completed, the prediction data is written to the external storage through the dram_pred_did and dram_pred_addrd interfaces, and the output signal o_inter_pred_done marks the end of the prediction.
[0094] The inter_pred_luma module is responsible for inter-frame prediction of the luminance component. Compared with the prediction of the chrominance component, the processing requirements of the luminance component are higher, so the module supports larger data widths and storage operations. The block position information is received through the i_xPb and i_yPb interfaces, and the motion vectors i_mvx and i_mvy are used to locate the prediction area of the reference frame. The module is designed with an architecture that supports multi-reference frame prediction, which can flexibly handle prediction requirements under different frame structures. The reference area range for luminance prediction is clearly defined by o_ref_start_x, o_ref_end_x, o_ref_start_y and o_ref_end_y. After the prediction is completed, it is indicated by the o_inter_pred_done signal, and the prediction result is stored in the external memory.
[0095] like Fig.11 The figure shows a schematic diagram of the motion estimation and compensation IP. The motion estimation and compensation IP includes two modes, Merge and AMVP, which are used to calculate motion vectors and perform motion compensation. The motion estimation IP is implemented, which receives the current frame macroblock and the reference frame and executes the motion estimation algorithm to find the best motion vector. The motion estimation module needs to calculate the pixel differences between multiple candidate blocks and select the best match.
[0096] Implement the motion compensation IP, which receives the reference frame, motion vector and current frame macroblock, and performs pixel block movement to generate an estimate. Ensure that the motion compensation module can correctly move the reference frame pixel block to the correct position.
[0097] Motion estimation and compensation are key steps in the H.265 / HEVC video decoding process to exploit the temporal redundancy between video frames. These steps allow the encoder to predict the position of macroblocks in the current frame based on previous reference frames, and then reduce the prediction error through compensation.
[0098] The goal of motion estimation is to find the position in the previous reference frame that best matches the macroblock of the current frame to generate a motion vector.
[0099] ①For each macroblock, the encoder searches for candidate blocks within a range and then calculates the difference between each candidate block and the current macroblock.
[0100] ② Motion estimation can use full search (Full Search) or other more efficient algorithms, such as block matching algorithm (Block Matching).
[0101] ③Finally, the motion estimation module outputs a motion vector, which indicates the position of the current macroblock in the reference frame.
[0102] The goal of motion compensation is to produce an estimate of a macroblock in the current frame by offsetting a block of pixels in a reference frame by a motion vector.
[0103] ① The motion compensation module receives the motion vector and the reference frame, and then moves the pixel block in the reference frame to the corresponding position in the current frame.
[0104] ②When the reference frame pixel block is in the correct position, the macroblock estimate of the current frame is constructed.
[0105] The motion estimation and compensation IP module calculates and applies the motion vector field (MVF) of the current macroblock through a complex hardware architecture, thereby efficiently completing the prediction and compensation operations of motion information.
[0106] The module flexibly supports different block partitioning modes and parallel computing configurations by receiving multiple input signals, such as the starting coordinates i_x0 and i_y0 of the current block, block size information i_nPbW and i_nPbH, and segmentation mode information i_part_mode. In addition, the input signal i_slice_temporal_mvp_enabled_flag indicates whether the temporal motion vector prediction (TMVP) is enabled, and i_num_ref_idx provides the number of reference frames. The reference frame information used by the module is passed in by interfaces such as i_ref_idx and i_delta_poc, so as to combine the motion vector of the reference frame for high-precision motion compensation.
[0107] In order to achieve the complexity of prediction, the module also supports storage management and direct memory access (DMA) functions. Interfaces such as m_axi_arvalid, m_axi_araddr and m_axi_rdata are responsible for interacting with external storage devices to achieve efficient loading and storage of reference data. By inputting the neighboring block motion vectors provided by signals such as i_left_mvf, i_up_mvf and i_left_up_mvf, the module combines the partition mode of the current block and the prediction mode i_predmode to complete the prediction calculation of the motion vector and output the final motion vector o_mvf.
[0108] When the temporal motion vector prediction (TMVP) function is turned on, the module can calculate the temporal position of the reference block according to i_cur_poc_diff and i_col_pic_poc, and complete the motion vector derivation across frames in combination with the reference frame number i_col_pic_dpb_slot. To improve the prediction accuracy, the module supports multi-reference frame selection and prediction based on fusion mode (Merge Mode), and input signals such as i_merge_flag and i_merge_idx realize the selection of fusion candidates.
[0109] The control and feedback of the internal state of the module are indicated by signals such as o_mv_done and o_col_param_fetch_done, ensuring the timing synchronization and operation reliability of each processing stage. Its design fully considers the efficient use of hardware resources and the hierarchical implementation of algorithm complexity, and can support the needs of real-time high-resolution video decoding. It is a core component of the H.265 / HEVC decoding system.
[0110] like Fig.12 The figure shows a schematic diagram of the filtering IP. The filtering IP includes an SAO filtering module and a deblocking filtering module. The deblocking filtering module receives the decoded image data, identifies the boundary pixels, and applies a filtering algorithm to reduce the blocking effect. The block size, boundary type, and pixel value need to be considered to determine the filtering method. The SAO filtering module receives the decoded image data and adjusts the pixel value based on the neighborhood information around the pixel to reduce artifacts. Different filtering methods are required to deal with edge artifacts and blockiness.
[0111] Deblocking filter and SAO (Sample Adaptive Offset) filter in the H.265 / HEVC video decoding process are important steps to improve decoding quality. These filtering processes are used to remove block effects and other artifacts caused by compression, thereby providing clearer and more accurate images. Deblocking filter can smooth the boundaries between adjacent blocks, making the image look more continuous and clear. SAO is divided into edge filtering and in-band filtering, the former is used to deal with edge artifacts, and the latter is used to deal with block effects.
[0112] The filter IP module is the core functional module for implementing deblocking filter (DF) and sample adaptive offset filter (SAO) in the H.265 / HEVC decoding process. Its design goal is to enhance image quality and reduce visual artifacts while maintaining the efficiency and real-time performance of hardware processing. The filter module contains two sub-modules, filter_32 and filter_64, which correspond to the filtering requirements of different block sizes respectively, and can adapt to various resolutions and performance requirements through flexible configuration.
[0113] The module performs independent filtering operations on the luminance (Luma) and chrominance (Chroma) components based on input signals such as the starting coordinates i_x0 and i_y0, block boundary information i_last_col, i_last_row and its width and height parameters i_last_col_width and i_last_row_height, and the QP values of intra-frame and inter-frame blocks, offset parameters i_qp_cb_offset, i_qp_cr_offset and SAO parameters i_sao_param, etc. The module uses a highly parallel filtering architecture to support simultaneous processing of block boundary signals i_bs_ver and i_bs_hor in the horizontal and vertical directions, ensuring that boundary strength judgment and pixel correction are completed in real-time decoding.
[0114] During the filtering process, the SAO module reduces the error introduced by quantization by counting and adjusting the pixel values at the block boundaries, while the deblocking filter module adjusts the pixel gradient according to the characteristics of the block boundaries to smooth the transition between blocks. The module also passes the fd_filter and fd_deblock feedback status information through registers to coordinate the overall filtering process. The module's direct memory access (DMA) interfaces such as m_axi_awaddr and m_axi_wdata are responsible for writing the filtering results, significantly improving data transmission efficiency.
[0115] The main difference between the filter_32 and filter_64 modules is reflected in the data processing width and cache design. Filter_64 further optimizes the processing capability of high-resolution video through a larger input data bandwidth and more on-chip caches (such as dram_rec_doa and dram_rec_dob). By configuring SAO and deblocking filter parameters and flexible hardware interfaces, the module can efficiently support complex filtering requirements in dynamic scenes and provide excellent image quality and computational efficiency for H.265 / HEVC video decoding.
[0116] Furthermore, the present invention also proposes a 4K video hard decoding acceleration method in which ARM and FPGA are coordinated. Using the 4K video hard decoding acceleration system in which ARM and FPGA are coordinated, the method comprises the following steps:
[0117] Step 1: The system is powered on, and the cortex-A9 processor in the ARM module initializes the hardware parameter configuration of the FPGA module;
[0118] Step 2: The ARM module loads the H.265 video data stored on the SD card to the DDR storage. The FPGA module accesses the DDR SDRAM memory chip of the ARM module through the AXI-Stream FIFO data flow transport IP, reads the H.265 image data, and writes it to the FPGA external DDR after data format and bit width conversion. The video stream decoding module starts working;
[0119] Step 3: The video decoding module reads the FPGA external DDR to obtain the video stream to be decoded, reads the NALU stream data through the NALU stream reading IP, manages the buffer of the stream through the NALU parsing IP and the NALU Buffer management IP, and parses the stream, respectively parsing the three parsing IPs of VPS, PPS, and SPS, and then uses the CABAC decoding IP to decode the three bins, and divides the decoded data into 3 paths, one path is sent to the dequantization IP and the inverse transformation IP for inverse transformation and quantization, the other path is sent to the intra-frame prediction IP, and the last path is sent to the inter-frame prediction IP and the motion compensation IP, and finally the output video image is obtained through the filtering IP;
[0120] Step 4: After obtaining the decoded video image, the decoded video image is sent to the DDR SDRAM memory chip of the ARM module for storage. After the processor reads the DDR end data, it writes the decoded YUV420 video image into the SD card for storage.
[0121] In step 2, the AXI-Lite of the AXI-Stream FIFO data stream transfer IP is used by the ARM to configure the video stream data transfer channel of the AXI-Stream, including the length of the data to be transferred and the occupancy of the FIFO.
[0122] The AXI_STR_TXD interface of the AXI-Stream FIFO data stream transfer IP is used to transfer the data transmitted from the ARM-side DDR and convert it into video stream data in AXI-Stream format.
[0123] The detailed software control process is as Fig.13 shown.
[0124] Combined with specific embodiments, the present invention is also illustrated by comparative experiments:
[0125] After testing that the processing effect of each IP of the FPGA is correct, the entire FPGA project is actually tested on the board. The bitstream generated by the FPGA project is burned into the FPGA board, and the AXI-Stream data stream transfer IP, PLL clock IP, AXI4-Stream Data Width Converter data bit width conversion IP, and custom H.265 decoding IP are configured with the help of XILINX's VIVADO SDK software. The H.265 decoder IP mainly includes NALU bitstream reading IP, NALU parsing IP, NALUBuffer management IP, CABAC decoding IP, inverse quantization IP, intra prediction IP, inter prediction IP, inverse transform IP, motion compensation IP, and filtering IP. The video decoding rates of 3840x2160 resolution under different CPU / FPGA configurations are compared as shown in the following table:
[0126]
[0127] It can be seen that the higher the clock frequency, the faster the FPGA processing speed, and the more obvious the comparison with the CPU processing speed; under the action of a 250MHz clock frequency, the decoding of 4K resolution video can reach about 30 frames per second, which is a great improvement in speed compared with CPU1.
[0128] In summary, the present invention has the following beneficial effects:
[0129] ① High performance: Compared with the pure software decoding scheme, the present invention improves the decoding efficiency and can support the real-time decoding of 4K resolution video streams.
[0130] ② Low power consumption: Combining the collaborative optimization of the FPGA and the ARM processor effectively reduces the power consumption during the decoding process.
[0131] ③Flexible expansion: Modular IP design facilitates adaptation to other video formats (such as H.264) or higher resolution video decoding requirements
[0132] ④Easy to implement: Based on the Zynq SoC platform, it makes full use of existing resources and is easy to implement hardware and engineering deployment.
[0133] What is disclosed above is only a preferred embodiment of the present invention, and it certainly cannot be used to limit the scope of rights of the present invention. Ordinary technicians in this field can understand that all or part of the processes of the above embodiment and equivalent changes made according to the claims of the present invention still fall within the scope of the invention.
Claims
1. A 4K video hard decoding acceleration system based on ARM and FPGA, characterized in that: It includes an ARM module, an FPGA module, an FPGA external DDR and an AXI bus transmission protocol, wherein the ARM module and the FPGA module are connected and communicated via the AXI bus transmission protocol, and the FPGA external DDR is separately connected to the FPGA; The ARM module includes a cortex-A9 processor, an AXI communication interface, a DDR storage and an SD card. The cortex-A9 processor is connected to the AXI bus transmission protocol through the AXI communication interface, and the cortex-A9 processor is also bidirectionally connected to the DDR storage and the SD card respectively; The FPGA module includes a video decoding module, a video timing control module, an AXI data bit width conversion module, an AXI-Stream FIFO data flow transport IP, and a video format conversion module. The video timing control module is respectively connected to the video decoding module and the AXI bus transmission protocol connection for timing synchronization control. The video decoding module obtains video stream data through the AXI data bit width conversion module and the AXI-Stream FIFO data flow transport IP in turn, and then outputs it through the video format conversion module. One end of the video format conversion module is bidirectionally connected to the FPGA external DDR, and the other end outputs unidirectionally to the AXI bus transmission protocol.
2. The 4K video hard decoding acceleration system of ARM and FPGA collaboration as claimed in claim 1, characterized in that: The video decoding module includes NALU code stream reading IP, NALU parsing IP, NALU Buffer management IP, CABAC decoding IP, inverse quantization and inverse transformation IP, intra-frame prediction IP, inter-frame prediction IP, motion estimation and compensation IP, and filtering IP; The NALU code stream reading IP is responsible for parsing the data content of the network abstraction layer unit and passing it to the subsequent decoding module for processing; The NALU Buffer management IP is responsible for the network abstraction layer unit data buffer management; The NALU parsing IP parses the data output by the NALU Buffer management IP, and according to the provisions of the H.265 standard, parses and outputs the syntax elements corresponding to the code stream; The CABAC decoding IP is responsible for decoding the input bitstream, passing the video bitstream to the CABAC decoding module via the data path, and passing the decoded symbols back to the main video decoding process; The inverse quantization and inverse transformation IP is responsible for receiving the quantization coefficient and the inverse quantization step as input, performing the inverse quantization operation, and restoring the compressed and encoded video data to the original pixel value; The intra prediction IP performs intra prediction on luminance and chrominance, receives the image block and the selected prediction mode as input, and generates prediction values; The inter-frame prediction IP performs inter-frame prediction on brightness and chrominance respectively, receives the current frame, reference frame and motion vector as input by means of data interaction with the inverse quantization and inverse transformation IP and the motion estimation and compensation IP, generates a predicted frame, adds the predicted frame and the residual decoded by the inverse transformation and inverse quantization module, obtains a decoded frame, and finally sends the decoded frame to the filtering module for filtering; The motion estimation and compensation IP is used to calculate motion vectors and perform motion compensation; The filtering IP includes a SAO filtering module and a deblocking filtering module. The deblocking filtering module receives the decoded image data, identifies boundary pixels, and applies a filtering algorithm to reduce the blocking effect; the SAO filtering module receives the decoded image data and adjusts the pixel value according to the neighborhood information around the pixel to reduce artifacts.
3. The 4K video hard decoding acceleration system of ARM and FPGA collaboration as claimed in claim 2, characterized in that: The AXI bus transmission protocol includes three different interface types: AXI-Lite, AXI4 and AXI-Stream.
4. The 4K video hard decoding acceleration system of ARM and FPGA collaboration as claimed in claim 3, characterized in that: The DDR storage includes a DDR controller and a DDR SDRAM memory chip, which is used for reading cache of the video stream to be decoded and local cache of the video stream after decoding.
5. The 4K video hard decoding acceleration system of ARM and FPGA collaboration as claimed in claim 4, characterized in that: The SD card is used to store the video stream to be decoded and the video stream after decoding, and the reading and writing of data are controlled by the cortex-A9 processor.
6. A 4K video hard decoding acceleration method using ARM and FPGA, using the 4K video hard decoding acceleration system using ARM and FPGA as described in any one of claims 1 to 5, characterized in that: The following steps are involved: Step 1: The system is powered on, and the cortex-A9 processor in the ARM module initializes the hardware parameter configuration of the FPGA module; Step 2: The ARM module loads the H.265 video data stored on the SD card to the DDR storage. The FPGA module accesses the DDR SDRAM memory chip of the ARM module through the AXI-Stream FIFO data flow transport IP, reads the H.265 image data, and writes it to the FPGA external DDR after data format and bit width conversion. The video stream decoding module starts working; Step 3: The video decoding module reads the FPGA external DDR to obtain the video stream to be decoded, reads the NALU stream data through the NALU stream reading IP, manages the buffer of the stream through the NALU parsing IP and the NALU Buffer management IP, and parses the stream, respectively parsing the three parsing IPs of VPS, PPS, and SPS, and then uses the CABAC decoding IP to decode the three bins, and divides the decoded data into 3 paths, one path is sent to the dequantization IP and the inverse transformation IP for inverse transformation and quantization, the other path is sent to the intra-frame prediction IP, and the last path is sent to the inter-frame prediction IP and the motion compensation IP, and finally the output video image is obtained through the filtering IP; Step 4: After obtaining the decoded video image, the decoded video image is sent to the DDRSDRAM memory chip of the ARM module for storage. After the processor reads the DDR end data, it writes the decoded YUV420 video image into the SD card for storage.
7. The 4K video hard decoding acceleration method of ARM and FPGA collaboration as claimed in claim 6, characterized in that: In step 2, the AXI-Lite of the AXI-Stream FIFO data flow handling IP is used to configure the video stream data transmission channel of the AXI-Stream on ARM, including the length of the data to be transmitted and the occupancy of the FIFO; The AXI_STR_TXD interface of the AXI-Stream FIFO data stream handling IP is used to handle the data transmitted from the ARM DDR and convert it into video stream data in AXI-Stream format.
8. The 4K video hard decoding acceleration method of ARM and FPGA collaboration as claimed in claim 7, characterized in that: VPS is the video parameter set. The VPS parsing IP module is used to extract parameter information describing the decoder buffer and video sequence characteristics in the video stream; PPS stands for picture parameter set, and the PPS parsing IP module focuses on parsing the picture parameter set, especially for fine-grained control of the slice layer; SPS is a sequence parameter set. The SPS parsing IP module further parses the sequence parameter set in the NALU code stream, including the video frame width, height, chroma format, and maximum encoding block size parameters.
Citation Information
Cited By
Cloud server audio and video data transmission method based on ARM architecture
CN121217918A
An audio and video data transmission method based on an ARM architecture cloud server
CN121217918B