A spaceborne heterogeneous H.264 video compression and encoding system and encoding method

By adopting the heterogeneous architecture of domestic CPU and GPU in the satellite system, combining parallel processing and new "Z-shaped" scanning sequence, the problems of timing resource limitation and slow processing speed in the satellite video compression technology are solved, and efficient video compression and independent and controllable domestic technology are achieved.

CN116527895BActive Publication Date: 2025-06-20NAT SPACE SCI CENT CAS
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310374018.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-10
Publication Date
2025-06-20
Estimated Expiration
2043-04-10

AI Technical Summary

Technical Problem

The existing satellite-based video compression technology faces the challenges of timing resource limitations, slow processing speed and independent and controllable domestic production, and it is difficult to effectively process massive video image data.

Method used

A satellite-borne heterogeneous H.264 video compression encoding system based on domestic CPUs and GPUs is proposed, using a satellite-borne heterogeneous parallel forward coding subsystem and a parallel accelerated backward reconstruction subsystem. Through macroblock-level parallel processing and a new "Z-shaped" scanning sequence, efficient compression of video images is achieved.

Benefits of technology

It effectively reduces the amount of computing, improves the real-time nature of video compression, has independent intellectual property rights, solves the problem of limited video compression bandwidth in aerospace applications, and provides a technical basis for video compression in orbit applications in subsequent models of tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116527895B_ABST
    Figure CN116527895B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of spaceborne heterogeneous video compression technology, and particularly relates to a spaceborne heterogeneous H.264 video compression encoding system and an encoding method. The system is implemented based on domestic CPU and GPU, and includes a spaceborne heterogeneous parallel forward encoding subsystem and a parallel acceleration backward reconstruction subsystem. Among them, the spaceborne heterogeneous parallel forward encoding subsystem is used to combine the reconstructed images output by the parallel acceleration backward reconstruction subsystem, and perform macroblock-level parallel intra-frame prediction and inter-frame prediction on the video image sequence obtained in real time by the spaceborne high-resolution imaging device frame by frame according to a new "Zigzag" scanning order. After macroblock-level parallel DCT transformation, macroblock-level parallel quantization and encoding, a compressed bitstream is obtained. The parallel acceleration backward reconstruction subsystem is used to perform macroblock-level parallel inverse quantization and macroblock-level parallel inverse DCT transformation on the coefficient data block after macroblock-level parallel quantization to obtain a residual data block, and then obtain a reconstructed image through macroblock-level parallel loop filtering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of spaceborne heterogeneous video compression technology, and particularly to a spaceborne heterogeneous H.264 video compression and encoding system and an encoding method. Background Art

[0002] With the increasing complexity and diversification of China's space missions, the quantity and precision of spaceborne imaging payloads have been continuously improved, and the volume of video image data has grown geometrically. Limited by the geographical distribution of ground receiving stations, satellites generally adopt a mechanism of storing first and then downlinking. That is, the collected data is first stored in the on-board memory, and then transmitted to the ground when passing over the station. Due to the limitations of the satellite in terms of volume, weight, and power consumption, the satellite data storage and the downlink transmission bandwidth are limited, and the massive video image data poses great pressure on satellite data management. Spaceborne video compression is the key technology to solve this problem. Using a high-performance heterogeneous system to process on-board tasks and on-board payload data is a current research hotspot.

[0003] As follows Figure 1 The following is a traditional H.264 encoding block diagram, and the entire process is generally implemented on an FPGA or a CPU. The system includes a forward encoding subsystem from left to right and a backward reconstruction subsystem from right to left; among them,

[0004] The forward encoding subsystem includes an intra prediction module, an inter prediction module, a transform module, a quantization module, and an entropy encoding module. The above modules are all controlled by the same controller, and are used to encode the input video image information to obtain a compressed bitstream.

[0005] The backward reconstruction subsystem includes an inverse quantization module, an inverse transform module, and a filtering module. The above modules are all controlled by the same controller, and are used to provide a reconstructed frame for the inter prediction module.

[0006] The existing compression algorithms mainly include the H.26X series, the MPEG series, the AVS series, and the HEVC. Among them, HEVC / H.265 / H.266 / AVS2 / AVS3 are mainly for high-definition and ultra-high-definition videos, with relatively high complexity; in the MPEG series, MPEG-1 and MPEG-2 have serious blocking effects and cannot be edited and played back. Compared with H.264 at the same bit rate, MPEG-4 has noise at the image edge and slow network transmission speed, that is, the MPEG series has the characteristics of poor image quality and slow speed; in comparison, H.264 has moderate complexity and is suitable for video transmission in a wireless channel with severe interference, and can obtain better image quality in real time, which is suitable for space applications.

[0007] Existing heterogeneous video compression technologies mainly focus on FPGA - based heterogeneity, DSP - based heterogeneity, and CPU + GPU - based heterogeneity. However, for FPGA - based heterogeneous implementation, various timing and FPGA resource limitations need to be considered, making it difficult to implement complex image - processing algorithms. For DSP - based heterogeneous implementation, the processing speed is relatively slow. For CPU + GPU - based heterogeneous implementation, current research is almost all based on the heterogeneity of Intel CPU + Nvidia GPU. Due to special aerospace applications, it is very necessary and urgent to develop China's independent heterogeneous technology. Summary of the Invention

[0008] Aiming at the problems in the above - mentioned technologies such as the H.264 serial encoder being restricted by timing resources, slow processing speed, and issues related to independent control and localization, the purpose of the present invention is to overcome the above - mentioned defects of the existing technologies and propose a space - borne heterogeneous H.264 video compression coding system and coding method.

[0009] To achieve the above purpose, the present invention proposes a space - borne heterogeneous H.264 video compression coding system. The system is based on domestic CPU and GPU and includes a space - borne heterogeneous parallel forward - coding subsystem and a parallel - acceleration backward - reconstruction subsystem. Among them,

[0010] The space - borne heterogeneous parallel forward - coding subsystem is used to combine the reconstructed images output by the parallel - acceleration backward - reconstruction subsystem, and perform macro - block - level parallel intra - prediction and inter - prediction on the video image sequence obtained in real - time by the space - borne high - resolution imaging device frame by frame according to the new "Zigzag" scanning order. After macro - block - level parallel DCT transformation, macro - block - level parallel quantization, and coding, the compressed bitstream is obtained.

[0011] The parallel - acceleration backward - reconstruction subsystem is used to perform macro - block - level parallel inverse quantization and macro - block - level parallel inverse DCT transformation on the coefficient data block after macro - block - level parallel quantization to obtain the residual data block, and then obtain the reconstructed image through macro - block - level parallel loop filtering.

[0012] As an improvement of the above system, the space - borne heterogeneous parallel forward - coding subsystem includes: a reading and parsing module and an entropy - coding module deployed on the CPU, and a macro - block - level parallel intra - prediction module, a macro - block - level parallel inter - prediction module, a macro - block - level parallel DCT transformation module, and a macro - block - level parallel quantization module deployed on the GPU. Among them,

[0013] The reading and parsing module is used to read the video image sequence, parse the information of all macro - blocks that can be processed in parallel frame by frame, save it to intermediate variables, and send it to the macro - block - level parallel intra - prediction module.

[0014] The macro-block level parallel intra prediction module is used to predict the pixels of the current block by adopting a new "Zigzag" scanning order and decoding adjacent pixels of the parallel acceleration backward reconstruction subsystem;

[0015] The macro-block level parallel inter prediction module is used to use the block most similar to the current block in adjacent reference frames reconstructed by the parallel acceleration backward reconstruction subsystem as a prediction block, and perform motion compensation according to the calculated motion vector to obtain a predicted image block;

[0016] The macro-block level parallel DCT transform module is used to redistribute the predicted residual data block in the transform domain through macro-block level parallel DCT transform to reduce the correlation between pixels; the predicted residual data block is obtained by subtracting the predicted image block from the original image block stored in the intermediate variable;

[0017] The macro-block level parallel quantization module is used to obtain a quantized coefficient data block through macro-block level parallel quantization processing to achieve data compression;

[0018] The entropy coding module is used to perform entropy coding on the quantized coefficient data block by combining the intra prediction information of the macro-block level parallel intra prediction module and the motion information of the macro-block level parallel inter prediction module, reduce the video coding information volume by removing information entropy redundancy, and output the code stream to the network abstraction layer.

[0019] As an improvement of the above system, the macro-block level parallel intra prediction module includes luminance 4×4 sub-blocks, and the sub-block numbers range from 0 to 15;

[0020] Adopt the scanning order corresponding to sub-block numbers 0->1->2->4->3->8->6->5->9->7->10->12->11->13->14->15.

[0021] As an improvement of the above system, the processing processes of the macro-block level parallel DCT transform module and the macro-block level parallel quantization module include: combining the DCT transform and quantization processes into one, on the basis of realizing by multiplication and shift and adopting integer operations, using GPU parallel processing for multiplication operations to improve the real-time performance of coding compression; by adjusting the quantization step QP value, performing coarse quantization on the high-frequency part and fine quantization on the low-frequency part to reduce visual redundancy and quantization error.

[0022] As an improvement of the above system, the parallel acceleration backward reconstruction subsystem includes a macro-block level parallel inverse quantization module, a macro-block level parallel inverse DCT module, and a macro-block level parallel loop filter module deployed on the GPU, where,

[0023] The macro-block level parallel inverse quantization module is used to obtain an inverse quantized coefficient data block by performing macro-block level parallel inverse quantization on the quantized coefficient data block;

[0024] The macro-block level parallel inverse DCT module is used to obtain a residual data block by performing macro-block level parallel inverse DCT transformation on the quantized coefficient data block;

[0025] The macro-block level parallel loop filtering module is used to perform loop filtering on the reconstructed data block to remove blocking artifacts and obtain a reconstructed pixel block, where the reconstructed data block is obtained by adding the residual data block and the predicted image block.

[0026] As an improvement to the above system, the macro-block level parallel loop filtering module is designed for parallel optimization according to the parameter characteristics of the decision filtering strength BS. Taking a 16×16 block as the calculation unit, an image has 8 sides and 128 boundary points. An image frame is given to 1 block and 128 threads are used for parallel processing.

[0027] As an improvement to the above system, the domestic CPU is Loongson CPU and the domestic GPU is Vigo GPU.

[0028] On the other hand, the present invention proposes a spaceborne heterogeneous H.264 video compression and coding method, which is implemented based on the above system. The method includes a forward coding process and a backward reconstruction process:

[0029] The reading and parsing module reads the video image sequence, parses the information of all macro-blocks that can be processed in parallel frame by frame, saves it to an intermediate variable, and sends it to the macro-block level parallel intra prediction module to enter the forward coding process:

[0030] The macro-block level parallel intra prediction module adopts a new "Zigzag" scanning order to predict the pixels of the current block by parallelly accelerating the decoded adjacent pixels of the backward reconstruction subsystem.

[0031] The macro-block level parallel inter prediction module uses the block most similar to the current block in the adjacent reference frames reconstructed by the backward reconstruction subsystem as the prediction block through parallel acceleration, and performs motion compensation according to the calculated motion vector to obtain a predicted image block;

[0032] The macro-block level parallel DCT transformation module redistributes the predicted residual data block in the transform domain through macro-block level parallel DCT transformation to reduce the correlation between pixels; the predicted residual data block is obtained by subtracting the predicted image block from the original image block saved in the intermediate variable;

[0033] The macro-block level parallel quantization module obtains the quantized coefficient data block through macro-block level parallel quantization processing to achieve data compression;

[0034] The entropy encoding module performs entropy encoding on the quantized coefficient data block by combining the intra-prediction information of the macroblock-level parallel intra-prediction module and the motion information of the macroblock-level parallel inter-prediction module, reduces the video coding information volume by removing information entropy redundancy, and outputs the bitstream to the network abstraction layer. It should be noted that H.264 is divided into the Video Coding Layer (VCL for short) and the Network Abstraction Layer (NAL for short) at the system level. The compressed and encoded video data (VCL data) is packed into NAL units and transmitted through the network.

[0035] The backward reconstruction process includes:

[0036] The macroblock-level parallel inverse quantization module obtains the inverse quantized coefficient data block by performing macroblock-level parallel inverse quantization on the quantized coefficient data block;

[0037] The macroblock-level parallel inverse DCT module obtains the residual data block by performing macroblock-level parallel inverse DCT transformation on the inverse quantized coefficient data block;

[0038] The macroblock-level parallel loop filter module performs loop filtering on the reconstructed data block to remove blocking artifacts and obtains the reconstructed pixel block, where the reconstructed data block is obtained by adding the residual data block and the predicted image block.

[0039] Compared with the prior art, the advantages of the present invention are as follows:

[0040] 1. It has the advantages of parallel acceleration, effectively reducing the computational amount, and having independent intellectual property rights. The present invention is conducive to solving the restrictions of aerospace developed countries on China's aerospace technology and high-performance aerospace-grade devices, ensuring China's aerospace security, reducing the overall project cost; and is conducive to the development of aerospace computer application technology with independent intellectual property rights in China;

[0041] 2. A new "Zigzag" scanning order is applied in the on-orbit video compression spaceborne heterogeneous parallel forward encoding subsystem of the present invention. This scanning order can reduce strong data correlations, reduce the waiting times required for 4×4 prediction in a macroblock, and is conducive to parallel implementation. The specific scanning order is: 0->1->2->4->3->8->6->5->9->7->10->12->11->13->14->15;

[0042] 3. The macroblock-level parallel loop filter module in the parallel acceleration backward reconstruction subsystem for on-orbit video compression of the present invention is optimized for parallel design according to the parameter characteristics of the decision filtering strength BS, and can accelerate the filtering speed with little impact on the image quality;

[0043] 4. The present invention is oriented towards the requirements of on-orbit video compression and the autonomous control of key devices, and can solve the contradiction between the downlink of a large amount of on-orbit video data and the limited satellite downlink transmission bandwidth. It has the advantages of being unrestricted by FPGA timing and resources and enabling parallel acceleration, and can provide a technical basis for the on-orbit application of video compression in subsequent model tasks. Description of the Drawings

[0044] Figure 1 is the flow block diagram of a traditional H.264 video compression and encoding system;

[0045] Figure 2 is the design block diagram of the H.264 video compression and encoding system based on domestic CPU + GPU on-board heterogeneous of the present invention;

[0046] Figure 3 is the existing "Zig-Zag" scan order; where (a) is the traditional intra-frame luminance 4×4 sub-block Zig-Zag scan order, and (b), (c), (d), (e), (f) are various existing improved scan orders;

[0047] Figure 4 is a new "Zig-Zag" scan order in the on-board heterogeneous parallel forward encoding subsystem for on-orbit video compression of the present invention;

[0048] Figure 5 is the flow chart of the macro-block level parallel DCT transform and macro-block level parallel quantization module in the on-board heterogeneous parallel forward encoding subsystem for on-orbit video compression of the present invention;

[0049] Figure 6 is the flow chart of the macro-block level parallel loop filtering module in the parallel acceleration backward reconstruction subsystem for on-orbit video compression of the present invention;

[0050] Figure 7 is the schematic diagram of the on-board heterogeneous H.264 video compression and encoding of the present invention. Detailed Embodiment

[0051] The technical solution of the present invention will be described in detail below with reference to the drawings and embodiments.

[0052] Embodiment 1

[0053] As Figure 2 shown, Embodiment 1 of the present invention proposes an on-board heterogeneous H.264 video compression and encoding system, which is implemented based on domestic CPU and GPU and carried on an on-orbit satellite, and includes an on-board heterogeneous parallel forward encoding subsystem and a parallel acceleration backward reconstruction subsystem; wherein,

[0054] The on - satellite heterogeneous parallel forward encoding subsystem is used to combine the reconstructed images output by the parallel acceleration backward reconstruction subsystem, and perform macro - block - level parallel intra - prediction and inter - prediction on each frame of the video image sequence obtained in real - time by the on - satellite high - resolution imaging device in accordance with a new "Zigzag" scanning order. After macro - block - level parallel DCT transformation, macro - block - level parallel quantization, and encoding, a compressed bitstream is obtained.

[0055] The parallel acceleration backward reconstruction subsystem is used to perform macro - block - level parallel inverse quantization and macro - block - level parallel inverse DCT transformation on the coefficient data block after macro - block - level parallel quantization to obtain a residual data block, and then obtain a reconstructed image through macro - block - level parallel loop filtering.

[0056] Among them, the on - satellite heterogeneous parallel forward encoding subsystem includes a reading and parsing module, a new "Zigzag" scanning order intra - prediction module, a macro - block - level parallel inter - prediction module, a macro - block - level parallel DCT transformation module, a macro - block - level parallel quantization module, and an entropy encoding module, which are used to perform on - orbit data encoding on the video image data obtained in real - time by the on - satellite high - resolution imaging device to obtain a compressed bitstream. Specifically,

[0057] The reading and parsing module is used to read the video image sequence, parse the information of all macro - blocks that can be processed in parallel frame - by - frame, save it to intermediate variables, and send it to the macro - block - level parallel intra - prediction module.

[0058] The macro - block - level parallel intra - prediction module is used to adopt the new "Zigzag" scanning order and predict the pixels of the current block through the adjacent pixels decoded by the parallel acceleration backward reconstruction subsystem.

[0059] The macro - block - level parallel inter - prediction module is used to use the block most similar to the current block in the adjacent reference frames reconstructed by the parallel acceleration backward reconstruction subsystem as the prediction block, perform motion compensation according to the calculated motion vector, and obtain a predicted image block.

[0060] The macro - block - level parallel DCT transformation module is used to redistribute the predicted residual data block in the transform domain through macro - block - level parallel DCT transformation to reduce the correlation between pixels; the predicted residual data block is obtained by subtracting the predicted image block from the original image block saved in the intermediate variable.

[0061] The macro - block - level parallel quantization module is used to obtain a quantized coefficient data block through macro - block - level parallel quantization processing to achieve data compression.

[0062] The entropy encoding module is used to perform entropy encoding on the quantized coefficient data block by combining the intra - prediction information of the macro - block - level parallel intra - prediction module and the motion information of the macro - block - level parallel inter - prediction module, reduce the video coding information volume by removing information entropy redundancy, and output the bitstream to the network abstraction layer.

[0063] Parallel acceleration backward reconstruction subsystem, including a macro-block level parallel inverse quantization module, a macro-block level parallel inverse DCT module, and a macro-block level parallel loop filter module deployed on the GPU, which are used to perform inverse quantization and inverse DCT transformation on the quantized coefficient data block to obtain a residual data block, and then obtain a reconstructed image through the loop filter module to provide a reference frame for the intra-frame / inter-frame prediction module. Specifically,

[0064] The macro-block level parallel inverse quantization module is used to perform macro-block level parallel inverse quantization on the quantized coefficient data block to obtain an inverse quantized coefficient data block;

[0065] The macro-block level parallel inverse DCT module is used to perform macro-block level parallel inverse DCT transformation on the inverse quantized coefficient data block to obtain a residual data block;

[0066] The macro-block level parallel loop filter module is used to perform loop filtering on the reconstructed data block to remove block artifacts and obtain a reconstructed pixel block, where the reconstructed data block is obtained by adding the residual data block and the predicted image block.

[0067] Figure 3 is the existing scan order. Among them,

[0068] (a) is the traditional intra-frame luminance 4×4 sub-block Zig-Zag scan order. Using this scan order will bring a large amount of data correlation, which is extremely unfavorable for parallel design. For example, when processing the 3rd sub-block, it is necessary to wait for the 2nd sub-block to complete reconstruction before prediction and mode discrimination can be performed, while at this time the 4th sub-block can be processed without correlation. If the scan method in Figure (a) is used, it will result in an extremely long intra-frame prediction path and waste a lot of unnecessary waiting time. Using this scan method requires 326 clock cycles.

[0069] (b), (c), (d), (e), and (f) are the scan orders proposed based on the path cost brought by quantitative analysis of the scan order that currently exist. Among them, (f) reduces the correlation to the minimum by removing modes 3 and 7 in some 4×4 blocks, but it will cause image quality loss. The number of strong correlations of (b), (c), (d), and (e) is the same. By adjusting the reference pixel modes of the blocks with weak correlation numbers, the same performance can be achieved. It only takes 163 clock cycles to complete a 4×4 sub-block in (d).

[0070] Figure 4 is a new "Zigzag" scan order in the on-orbit video compression spaceborne heterogeneous parallel forward encoding subsystem of the present invention. The number of strong correlations is 2, and the number of times to wait for a 4×4 prediction to be completed is the same as Figure 3The same as (b), (c), (d), and (e). Strong correlation means that the next block can only be processed after the reconstruction of the previous block is completed, while the reason for the weak correlation is the correlation generated by several modes of the reference pixels of the current 4×4 sub-block. The weak correlation involved in the present invention is generated by mode 3, mode 7, and mode 8. That is, when predicting mode 3, mode 7, and mode 8, the reference pixel point needs to use the reconstructed pixel value of the upper right or lower left 4×4 sub-block of the current 4×4 sub-block, and the reconstructed pixel may be in the processing process. The weak correlation in these methods can be eliminated by reordering the mode order, such as processing mode 3, mode 7, and mode 8 of the weakly correlated blocks last. For example, Figure 4 The 4×4 sub-blocks with weak correlation are numbered 2, 3, 5, 8, 11, and 14. These sub-blocks still have 6 modes available for prediction and calculation before processing modes 3, 7, and 8. As long as the upper right 4×4 reference pixel of the current 4×4 sub-block can be reconstructed before these 6 modes are processed, the above weak correlation can be eliminated.

[0071] Table 1 Correlation analysis of various scanning methods

[0072] Scanning method Number of strong correlations Number of weak correlations Method a 8 4 Method b 2 4 Method c 2 3 Method d 2 7 Method e 2 5 Method f 2 0 New method 2 6

[0073] Figure 5 This is a flow chart of the macroblock-level parallel DCT transform and macroblock-level parallel quantization module in the satellite-borne heterogeneous parallel forward coding subsystem for on-orbit video compression of the present invention. In addition to DCT transform, traditional transforms also include KL transform, but its complexity is relatively high, so it is not considered for use. Normally, H.264 encoding combines quantization division and transform normalization into one and implements them through multiplication and shift. Currently, there are also designs that use butterfly operations to reduce the amount of calculation in the process. The present invention not only combines quantization division and transform normalization into one and implements them through multiplication and shift, but also implements butterfly operations in a parallel manner, greatly reducing the amount of calculation in the process. Since the amount of multiplication calculation is large, GPU parallel acceleration is used.

[0074] Figure 6 This is a flow chart of the macroblock-level parallel loop filter module in the parallel accelerated backward reconstruction subsystem for on-orbit video compression described in the present invention. Parallel optimization design is performed based on the parameter characteristics that determine the filter strength BS. The process uses 16×16 blocks as calculation units, the number of edges of a frame of image is 8, and the number of boundary points is 128. The parameter values ​​required for boundary point calculation have been stored in the global memory. A frame of image can be directly handed over to 1 block, and 128 threads are used for parallel processing. The filtering speed is accelerated without affecting the image quality.

[0075] Example 2

[0076] likeFigure 7 As shown in Figure 7 , Embodiment 2 of the present invention proposes a spaceborne heterogeneous H.264 video compression and encoding method, including:

[0077] First, read the video sequence to be encoded on the Loongson CPU and enter the forward encoding process. Then, parse the information of all macroblocks that can be processed in parallel in a frame of image and save it to an intermediate variable. Then, transfer the data to the VIV VPU for macroblock-level parallel intra / inter-frame prediction, DCT transformation, and quantization processing. The quantized data will be transmitted back to the Loongson CPU for entropy encoding processing. The backward reconstruction loop includes an inverse quantization module, an inverse DCT transformation module, and a loop filtering module. The reconstructed image after loop filtering will be used as a reference frame to continue participating in the processing of subsequent frames. The data after encoding processing will be output to the network abstraction layer and transmitted to the ground receiving station through the data transmission system when the satellite passes by. The on-orbit real-time video compression function is realized.

[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the embodiments, those of ordinary skill in the art should understand that any modification or equivalent replacement of the technical solutions of the present invention does not depart from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A spaceborne heterogeneous H.264 video compression and encoding system, characterized in that The system is implemented based on domestic CPUs and GPUs, and includes an on-board heterogeneous parallel forward encoding subsystem and a parallel acceleration backward reconstruction subsystem. Among them, The on-board heterogeneous parallel forward encoding subsystem is used to combine the reconstructed images output by the parallel acceleration backward reconstruction subsystem, and perform macroblock-level parallel intra-frame prediction and inter-frame prediction on each frame of the video image sequence obtained in real time by the on-board high-resolution imaging device. After macroblock-level parallel DCT transformation, macroblock-level parallel quantization, and entropy encoding, a compressed bitstream is obtained. The parallel acceleration backward reconstruction subsystem is used to perform macroblock-level parallel inverse quantization and macroblock-level parallel inverse DCT transformation on the coefficient data blocks after macroblock-level parallel quantization to obtain residual data blocks, and then obtain reconstructed images through macroblock-level parallel loop filtering. The macroblock includes luminance 4×4 sub-blocks, and the sub-block numbers range from 0 to 15. The macroblock-level parallel intra-frame prediction adopts a scanning order corresponding to sub-block numbers 0->1->2->4->3->8->6->5->9->7->10->12->11->13->14->15.

2. The spaceborne heterogeneous H.264 video compression and encoding system according to claim 1, characterized in that The on-board heterogeneous parallel forward encoding subsystem includes: a reading and parsing module and an entropy encoding module deployed on the CPU, and a macroblock-level parallel intra-frame prediction module, a macroblock-level parallel inter-frame prediction module, a macroblock-level parallel DCT transformation module, and a macroblock-level parallel quantization module deployed on the GPU. Among them, The reading and parsing module is used to read the video image sequence, parse the information of all macroblocks that can be processed in parallel frame by frame, save it to intermediate variables, and send it to the macroblock-level parallel intra-frame prediction module. The macroblock-level parallel intra-frame prediction module is used to predict the pixels of the current block through the adjacent pixels decoded by the parallel acceleration backward reconstruction subsystem. The macroblock-level parallel inter-frame prediction module is used to use the block most similar to the current block in the adjacent reference frames reconstructed by the parallel acceleration backward reconstruction subsystem as the prediction block, perform motion compensation according to the calculated motion vector, and obtain the predicted image block. The macroblock-level parallel DCT transformation module is used to redistribute the predicted residual data blocks in the transform domain through macroblock-level parallel DCT transformation to reduce the correlation between pixels. The predicted residual data blocks are obtained by subtracting the predicted image blocks from the original image blocks saved in the intermediate variables. The macroblock-level parallel quantization module is used to obtain quantized coefficient data blocks through macroblock-level parallel quantization processing to achieve data compression. The entropy encoding module is used to perform entropy encoding on the quantized coefficient data blocks by combining the intra-frame prediction information of the macroblock-level parallel intra-frame prediction module and the motion information of the macroblock-level parallel inter-frame prediction module, reduce the video coding information volume by removing information entropy redundancy, and output the bitstream to the network abstraction layer.

3. The spaceborne heterogeneous H.264 video compression and encoding system according to claim 2, characterized in that The processing procedures of the macroblock-level parallel DCT transform module and the macroblock-level parallel quantization module include: combining the DCT transform and quantization processes into one, and on the basis of implementing them through multiplication and shifting and using integer operations, utilizing the GPU to parallelize the multiplication operation to improve the real-time performance of encoding compression; by adjusting the quantization step QP value, coarsely quantizing the high-frequency part and finely quantizing the low-frequency part to reduce visual redundancy and quantization error.

4. The spaceborne heterogeneous H.264 video compression and encoding system according to claim 2, characterized in that The parallel accelerated backward reconstruction subsystem includes a macroblock-level parallel inverse quantization module, a macroblock-level parallel inverse DCT module, and a macroblock-level parallel loop filtering module deployed on the GPU. Among them, The macroblock-level parallel inverse quantization module is used to obtain the inverse quantized coefficient data block by performing macroblock-level parallel inverse quantization on the quantized coefficient data block; The macroblock-level parallel inverse DCT module is used to obtain the residual data block by performing macroblock-level parallel inverse DCT transform on the inverse quantized coefficient data block; The macroblock-level parallel loop filtering module is used to perform loop filtering processing on the reconstructed data block to remove block artifacts and obtain the reconstructed pixel block, and the reconstructed data block is obtained by adding the residual data block and the predicted image block.

5. The spaceborne heterogeneous H.264 video compression and encoding system according to claim 4, characterized in that The macroblock-level parallel loop filtering module takes a 16×16 block as the calculation unit. The number of sides of a frame of image is 8, and the number of boundary points is 128. One frame of image is given to 1 block, and 128 threads are used for parallel processing.

6. The spaceborne heterogeneous H.264 video compression and encoding system according to claim 1, characterized in that The domestic CPU is Loongson CPU, and the domestic GPU is Vigo GPU.

7. A spaceborne heterogeneous H.264 video compression and encoding method, implemented based on the system according to claim 3, the method comprising a forward encoding process and a backward reconstruction process: The reading and parsing module reads the video image sequence, parses the information of all macroblocks that can be processed in parallel frame by frame, saves it to intermediate variables, and sends it to the macroblock-level parallel intra prediction module to enter the forward encoding process: The macroblock-level parallel intra prediction module predicts the pixels of the current block by parallelly accelerating the decoded adjacent pixels of the backward reconstruction subsystem; The macroblock-level parallel inter-frame prediction module uses the block most similar to the current block in the adjacent reference frames reconstructed by the parallel accelerated backward reconstruction subsystem as the prediction block, and performs motion compensation according to the calculated motion vector to obtain the predicted image block; The macroblock-level parallel DCT transform module redistributes the predicted residual data block in the transform domain through macroblock-level parallel DCT transform to reduce the correlation between pixels; The predicted residual data block is obtained by subtracting the predicted image block from the original image block stored in the intermediate variable; The macroblock-level parallel quantization module obtains the quantized coefficient data block through macroblock-level parallel quantization processing to achieve data compression; The entropy encoding module performs entropy encoding on the quantized coefficient data block by combining the intra-frame prediction information of the macroblock-level parallel intra-frame prediction module and the motion information of the macroblock-level parallel inter-frame prediction module, reduces the video coding information volume by removing information entropy redundancy, and outputs the bitstream to the network abstraction layer; The backward reconstruction process includes: The macroblock-level parallel inverse quantization module obtains the inverse quantized coefficient data block by performing macroblock-level parallel inverse quantization on the quantized coefficient data block; The macroblock-level parallel inverse DCT module obtains the residual data block by performing macroblock-level parallel inverse DCT transform on the inverse quantized coefficient data block; The macroblock-level parallel loop filtering module performs loop filtering processing on the reconstructed data block to remove block artifacts and obtain the reconstructed pixel block, and the reconstructed data block is obtained by adding the residual data block and the predicted image block.

Citation Information

Patent Citations

  • Parallel processing method for implementing entropy coding link in HEVC based on CPU+GPU heterogeneous platform

    CN109391816A

  • Encoding method and decoding method for high sharpness video super strong compression

    CN1784008A

  • Device for encoding / decoding motion image, method therefor and recording medium storing a program to implement thereof

    KR1020120041502A