Method and Decoder for HDR Vivid Video Post-Processing
Through the video decoding post-processing method of CPU+GPU hybrid architecture, the problem of real-time decoding of HDR Vivid video in ultra-high-definition applications is solved, real-time monitoring and multi-screen adaptability of low-cost 8K ultra-high-definition HDR Vivid videos are achieved.
Patent Information
- Application Number
- CN202211229905.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-09
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-10-09
AI Technical Summary
The prior art cannot effectively realize ultra-high-definition real-time decoding monitoring of high dynamic range (HDR Vivid) videos, especially in 4K/8K applications, the terminal device cannot play in real time, and cannot adapt to the display differences of multiple screen parameters.
Adopting a hybrid CPU+GPU architecture, by dividing video data into multiple slices, parallel processing, establishing a GPU resource pool, and implementing parallel pipeline operations of HDR Vivid video decoding and post-processing, including H2D copy, CU general computing and D2H copying, optimizing hardware resource utilization.
Real-time decoding and monitoring of low-cost 8K ultra-high-definition HDR Vivid videos is realized, which improves real-time processing and system parallelism, and adapts to the display needs of multiple screen parameters.
Smart Images

Figure CN115665421B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of audio and video coding and decoding, and particularly relates to a method and a decoder for post-processing HDR Vivid video decoding. Background Art
[0002] In a high-dynamic range (HDR) Vivid ultra-high definition end-to-end system, a reference decoder is required to monitor the output bitstream to monitor the actual terminal playback effect of the HDR Vivid bitstream. The difference between the playback monitoring of the HDR Vivid bitstream and a traditional professional decoder is that the decoder needs to perform HDR Vivid post-processing on the decoded image according to the brightness range of the target display screen to correctly display the actual terminal playback effect of the HDR Vivid bitstream.
[0003] Generally, the bitstream of the program production end can be played on a terminal that has passed the HDR Vivid certification as a reference display effect. However, since each terminal screen has only one set of screen parameters and can only simulate and display one screen display effect, it cannot cover a vast number of terminals. On the other hand, as a reference decoder, the front-end monitoring device requires a relatively high playback processing restoration degree and display accuracy, and various links that may cause display errors need to be excluded. Due to the individual differences between the screen and the device, it cannot be guaranteed whether the current effect of the playback terminal is consistent with the post-processing effect of the original HDR Vivid. Finally, HDR Vivid post-processing is an image tone mapping technology that needs to perform tone mapping processing on each frame of image in real time according to the metadata in the bitstream, and the calculation amount is relatively large. Especially in 4K / 8K ultra-high definition applications, some terminals cannot achieve real-time playback.
[0004] The present invention realizes a method and a device for HDR Vivid decoding post-processing. Based on a CPU+GPU hybrid architecture, the decoding process and the HDR Vivid post-processing process are fully optimized and modularly parallelized and split, realizing a high-precision restoration of the HDR Vivid bitstream effect. Through reasonable device hardware configuration, real-time playback monitoring of ultra-high definition 4K / 8K HDR Vivid post-processing is achieved. By inputting different target screen parameters, the monitoring of the reference post-processing display effects of various screens can be realized, filling the gap in the industry, and promoting the popularization of the HDR Vivid standard and reducing the development and verification costs. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a method and a decoder for HDR Vivid video decoding post-processing, which can realize real-time decoding monitoring of 8K ultra-high definition HDR Vivid post-processing at a relatively low cost.
[0006] To solve the above technical problems, the present invention adopts the following technical solutions:
[0007] An embodiment of the present invention provides a method for post-processing HDR Vivid video decoding, including the following steps:
[0008] Divide the YUV data to be post-processed into N Slices, where N is set according to the device capabilities and image size;
[0009] Create and apply for an output frame V210 storage memory block;
[0010] Abstract the GPU resources into different capability types, including H2D copy, CU general computing, and D2H copy. The H2D copy is used to copy YUV data from memory to video memory; the CU general computing is used for Vivid post-processing calculation; the D2H copy is used to copy V210 data from video memory to memory. Establish a computing capability resource pool for each capability, and respectively establish an idle resource list and a used resource list to form a total GPU capability resource pool. Externally, apply for idle resources to execute corresponding operations, otherwise wait for resource release;
[0011] Create an independent thread module to perform HDR Vivid mapping curve parameter generation calculation;
[0012] Apply for H2D copy capability resources and execute the operation of copying YUV data from memory to video memory;
[0013] Obtain the HDR Vivid curve parameter generation result of the current frame from the curve parameter generation module.
[0014] Apply for CU general computing capability resources and execute the core calculation process of Vivid post-processing;
[0015] Apply for D2H copy capability and execute the operation of copying V210 data from video memory to the specified memory;
[0016] Perform synchronous SDI output processing on multiple Slices of the same frame of video.
[0017] Preferably, the division of Slices is achieved by calculating the start address and size of each YUV component of each Slice.
[0018] Preferably, the division of Slices is achieved by calculating the start address and size of each YUV component of each Slice, including:
[0019] During initialization, based on the image size and the YUV color space format, calculate the total sizes Y_len, U_len, and V_len of each YUV component data. According to the total sizes, calculate the slice block sizes for each component: Y_chunk = Y_len / N, U_chunk = U_len / N, and V_chunk = V_len / N, and save the above data. Before processing each frame of the image, obtain the starting addresses of the Y, U, and V components. Based on the starting addresses of each component and the slice block sizes calculated during initialization, calculate the starting addresses and block sizes of each slice for each YUV component.
[0020] Preferably, V210 is in a packed storage format. A frame of the image is represented by a single continuous block of memory. By creating the starting address of the V210 data block and the image size, calculate the target starting address of each slice.
[0021] Preferably, the external performs corresponding operations by applying for idle resources, otherwise waiting for resource release, including: for each slice of the image in the application scheduling layer, serially execute three sub-computation processes: H2D copy, Vivid post-processing calculation, and D2H copy in sequence. For each process executed, obtain computing resources from the corresponding GPU idle resource queue for processing, and at the same time move the computing resources in the idle queue to the used resource queue. If the idle queue is empty, wait for the used resources to be released back to the idle queue before performing the above operations. After the sub-computation process of each slice is completed, release the corresponding computing resources back to the idle queue; inside each slice, the three sub-computation steps are executed serially, and between multiple slices, they are executed in parallel, sharing computing resources.
[0022] Preferably, the creation of an independent thread module for generating HDR Vivid mapping curve parameters includes:
[0023] Create an independent thread module. After parsing the corresponding metadata during the video bitstream parsing process, hand it over to the thread module for processing. The thread module calculates the HDR Vivid mapping curve parameters and stores the output results according to the frame number. Finally, in the HDR Vivid post-processing calculation process, obtain the curve parameter calculation results of the current frame from this thread module to achieve asynchronous parallel execution of the curve calculation process and the video decoding, image format conversion, and electro-optical conversion processes.
[0024] Another aspect of the embodiments of the present invention provides a decoder for performing the HDR Vivid video decoding post-processing method described above, implemented based on a hybrid architecture of a server CPU and GPU.
[0025] Preferably, the YUV data to be post-processed is divided into N slices; an output frame V210 storage memory block is created and applied for; an independent thread module is created to perform the calculation of generating HDR Vivid mapping curve parameters; the synchronous SDI output processing for multiple slices of the same frame of video is performed in the CPU.
[0026] Preferably, the GPU resources are abstracted into different capability types, including H2D copy, CU general computing, and D2H copy. The H2D copy is used to copy YUV data from memory to video memory; the CU general computing is used for Vivid post-processing calculation; the D2H copy is used to copy V210 data from video memory to memory; a computing capability resource pool is established for each capability, an idle resource list and a used resource list are established respectively to form a total GPU capability resource pool, and externally, the corresponding operations are executed by applying for idle resources, otherwise wait for the resources to be released; apply for H2D copy capability resources and perform the operation of copying YUV data from memory to video memory; obtain the HDR Vivid curve parameter generation result of the current frame from the curve parameter generation module; apply for CU general computing capability resources and perform the core calculation process of Vivid post-processing; apply for D2H copy capability and perform the operation of copying V210 data from video memory to the specified memory by the GPU.
[0027] The present invention has the following beneficial effects: Through the CPU + multi-GPU heterogeneous solution, by forming a GPU resource pool and reasonably designing the parallel processes to form a pipeline operation, the present invention can achieve a high resource utilization rate, improve the parallelism and processing real-time performance of the system, and through reasonable hardware configuration, it can achieve real-time decoding and monitoring of 8K ultra-high-definition HDR Vivid post-processing at a relatively low cost. Description of the Drawings
[0028] Figure 1 It is a step flowchart of the method for HDR Vivid video decoding and post-processing according to an embodiment of the present invention. Detailed Embodiments
[0029] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0030] Refer to Figure 1 , which shows a step flowchart of the method for HDR Vivid video decoding and post-processing according to an embodiment of the present invention, including the following steps:
[0031] Divide the YUV data to be post-processed into N slices, where N is set according to the device capabilities and image size;
[0032] Create and apply for a storage memory block for the output frame V210;
[0033] Abstract the GPU resources into different capability types, including H2D (Host To Device) copy, CU (ComputeUnit) general computing, and D2H (Device To Host) copy. The H2D copy is used to copy YUV data from memory to video memory; CU general computing is used for Vivid post-processing calculations; the D2H copy is used to copy V210 data from video memory to memory. Establish a computing capability resource pool for each capability, and respectively establish an idle resource list and a used resource list to form a total GPU capability resource pool. Externally, apply for idle resources to perform corresponding operations, otherwise wait for resource release;
[0034] Create an independent thread module to perform HDR Vivid mapping curve parameter generation calculations;
[0035] Apply for H2D copy capability resources and perform the operation of copying YUV data from memory to video memory;
[0036] Obtain the HDR Vivid curve parameter generation result of the current frame from the curve parameter generation module.
[0037] Apply for CU general computing capability resources and perform the core calculation process of Vivid post-processing;
[0038] Apply for D2H copy capability and perform the operation of copying V210 data from video memory to the specified memory;
[0039] Perform synchronous SDI output processing on multiple slices of the same frame of video.
[0040] In the above process, except for the curve parameter coefficient generation process, other processes are independently processed for each pixel, and pixel-level parallel processing can be achieved, including curve parameter coefficient generation and HDR Vivid post-processing image processing. The two processes can be processed in parallel, and the curve parameter coefficient generation only needs to be completed before the compression ratio calculation of the HDR Vivid post-processing.
[0041] In an embodiment of the present invention, the division of slices is achieved by calculating the start address and size of each YUV component of each slice, including:
[0042] During initialization, based on the image size and the YUV color space format, calculate the total sizes Y_len, U_len, and V_len of each YUV component data. According to the total sizes, calculate the slice block sizes for each component: Y_chunk = Y_len / N, U_chunk = U_len / N, and V_chunk = V_len / N, and save the above data. Before processing each frame of the image, obtain the starting addresses of the Y, U, and V components. Based on the starting addresses of each component and the slice block sizes calculated during initialization, calculate the starting addresses and block sizes of each slice for each YUV component.
[0043] In an embodiment of the present invention, the external applies for idle resources to perform corresponding operations, otherwise waits for the release of resources, including: for each slice of the image in the application scheduling layer, sequentially executes three sub-computation processes of H2D copy, Vivid post-processing calculation, and D2H copy in series. For each process executed, obtain computing resources from the corresponding GPU idle resource queue for processing, and at the same time move the computing resources in the idle queue to the used resource queue. If the idle queue is empty, wait for the used resources to be released back to the idle queue and then execute the above operations. After the sub-computation process of each slice is completed, release the corresponding computing resources back to the idle queue; inside each slice, the three sub-computation steps are executed in series, and between multiple slices, they are executed in parallel, sharing computing resources.
[0044] In an embodiment of the present invention, creating an independent thread module for generating HDR Vivid mapping curve parameters includes:
[0045] Create an independent thread module. After parsing the corresponding metadata during the video bitstream parsing process, hand it over to the thread module for processing. The thread module calculates the HDR Vivid mapping curve parameters and stores the output results according to the frame number. Finally, in the HDR Vivid post-processing calculation process, obtain the curve parameter calculation results of the current frame from this thread module to achieve asynchronous parallel execution of the curve calculation process and the video decoding, image format conversion, and electro-optical conversion processes.
[0046] A decoder according to another embodiment of the present invention is used to perform the HDR Vivid video decoding post-processing method described above, and is implemented based on a hybrid architecture of a server CPU and GPU. Among them, the YUV data to be post-processed is divided into N slices; a storage memory block for the output frame V210 is created and applied for; an independent thread module is created to perform the calculation of generating HDR Vivid mapping curve parameters; the synchronous SDI output processing for multiple slices of the same frame video is performed in the CPU. The GPU resources are abstracted into different capability types, including H2D copy, CU general computing, and D2H copy. The H2D copy is used to copy YUV data from memory to video memory; the CU general computing is used for Vivid post-processing calculation; the D2H copy is used to copy V210 data from video memory to memory; a computing capability resource pool is established for each capability, an idle resource list and a used resource list are established respectively, forming a total GPU capability resource pool. Externally, the corresponding operation is executed by applying for idle resources. Otherwise, wait for the resource to be released; apply for the H2D copy capability resource and perform the operation of copying YUV data from memory to video memory; obtain the HDR Vivid curve parameter generation result of the current frame from the curve parameter generation module. Apply for the CU general computing capability resource and execute the core computing process of Vivid post-processing; apply for the D2H copy capability and execute the operation of copying V210 data from video memory to the specified memory, which is completed by the GPU.
[0047] Based on the CPU+GPU hybrid architecture, the GPU capabilities are abstracted into three capability resource pools, namely H2D copy, CU general computing, and D2H copy. The setting of the quantity of each resource can be reasonably configured based on the specific number of GPUs and the actual capabilities. Externally, the execution of specific computing modules is realized by applying for the corresponding idle resources of the GPU. By splitting the image to be processed into N slices for parallel processing, an operation pipeline is formed, which can maximize parallel processing while fully utilizing the hardware resources and reduce the overall latency.
[0048] It should be understood that the exemplary embodiments described herein are illustrative rather than restrictive. Although one or more embodiments of the present invention have been described in conjunction with the accompanying drawings, those of ordinary skill in the art should understand that various changes in form and detail can be made without departing from the spirit and scope of the present invention as defined by the appended claims.
Claims
1. A method for post-processing of decoded HDR Vivid video, characterized in that, The steps include: Dividing the YUV data to be post-processed into N Slices, where N is set according to the device capabilities and image size; Creating and applying for a storage memory block for the output frame V210; V210 is in a packed storage format, and one frame of the image is represented by a whole continuous memory. By creating the starting address and image size of the V210 data block, the target starting address of each Slice is calculated; Abstracting the GPU resources into different capability types, including H2D copy, CU general computing, and D2H copy. The H2D copy is used to copy YUV data from memory to video memory; the CU general computing is used for Vivid post-processing calculations; the D2H copy is used to copy V210 data from video memory to memory. A computing capability resource pool is established for each capability, and an idle resource list and a used resource list are established respectively to form a total GPU capability resource pool. Externally, the corresponding operations are executed by applying for idle resources, otherwise waiting for the resources to be released; Creating an independent thread module to perform the calculation for generating HDR Vivid mapping curve parameters; Applying for H2D copy capability resources and performing the operation of copying YUV data from memory to video memory; Obtaining the HDR Vivid curve parameter generation result of the current frame from the curve parameter generation module; Applying for CU general computing capability resources and performing the core calculation process of Vivid post-processing; Applying for D2H copy capability and performing the operation of copying V210 data from video memory to the specified memory; Performing synchronous SDI output processing on multiple Slices of the same frame of video.
2. The method for post - processing of decoded HDR Vivid video according to claim 1, wherein, The division of Slices is achieved by calculating the starting address and size of each YUV component of each Slice.
3. The method for post - processing of HDR Vivid video decoding according to claim 2, wherein, The division of Slices is achieved by calculating the starting address and size of each YUV component of each Slice, including: During initialization, according to the image size and YUV color space format, calculate the total size Y_len, U_len, and V_len of each YUV component data. According to the total size, calculate the Slice block size of each component: Y_chunk = Y_len / N, U_chuk = U_len / N, V_chunk = V_len / N, and save the above data. Before processing each frame of the image, obtain the starting address of the Y, U, and V components. According to the starting address of each component and the Slice block size calculated during initialization, calculate the starting address and block size of each YUV component of each Slice.
4. The method for post-processing of HDR Vivid video decoding according to claim 1, wherein, The external performs corresponding operations by applying for idle resources, otherwise waits for resource release, which includes: for each Slice of the image in the application scheduling layer, the three sub-computation processes of H2D copy, Vivid post-processing calculation, and D2H copy are serially executed in sequence. For each process executed, computing resources are obtained from the corresponding GPU idle resource queue for processing, and at the same time, the computing resources in the idle queue are moved into the used resource queue. If the idle queue is empty, wait for the used resources to be released back to the idle queue and then execute the above operations. After the sub-computation process of each Slice is completed, the corresponding computing resources are released back to the idle queue; inside each Slice, the three sub-computation steps are serially executed, and among multiple Slices, they are executed in parallel and share computing resources.
5. The HDR Vivid video decoding post-processing method according to claim 1, characterized in that, The creation of an independent thread module for HDR Vivid mapping curve parameter generation calculation includes: Create an independent thread module. After parsing the corresponding metadata during the video bitstream parsing process, hand it over to the thread module for processing. The thread module calculates the HDR Vivid mapping curve parameters and stores the output results according to the frame number. Finally, in the HDR Vivid post-processing calculation process, obtain the curve parameter calculation results of the current frame from this thread module to achieve asynchronous parallel execution of the processes including the curve calculation process, video decoding, image format conversion, and electro-optical conversion process.
6. A decoder, characterized in that, The method for performing HDR Vivid video decoding post-processing according to any one of claims 1 to 5 above is implemented based on a hybrid architecture of a server CPU and GPU.
7. The decoder according to claim 6, characterized in that Divide the YUV data to be post-processed into N Slices; create and apply for an output frame V210 storage memory block; create an independent thread module for HDR Vivid mapping curve parameter generation calculation; perform synchronous SDI output processing for multiple Slices of the same frame video in the CPU.
8. The decoder according to claim 6, characterized in that, Abstract the GPU resources into different capability types, including H2D copy, CU general computing, and D2H copy. The H2D copy is used to copy YUV data from memory to video memory; the CU general computing is used for Vivid post-processing calculation; the D2H copy is used to copy V210 data from video memory to memory; establish a computing capability resource pool for each capability, respectively establish an idle resource list and a used resource list to form a total GPU capability resource pool. The external performs corresponding operations by applying for idle resources, otherwise waits for resource release; apply for H2D copy capability resources and perform the operation of copying YUV data from memory to video memory; obtain the HDR Vivid curve parameter generation results of the current frame from the curve parameter generation module; apply for CU general computing capability resources and perform the core Vivid post-processing calculation process; apply for D2H copy capability and perform the operation of copying V210 data from video memory to the specified memory by the GPU.
Citation Information
Patent Citations
Remote-sensing image decompression method based on GPU (Graphics Processing Unit)
CN102158694A
Method and device for improving intelligent analysis performance
CN108206937A