Method for parallel decoding and rendering of sequence of JPEG images

By adopting CPU concurrent programming, thread allocation strategy, isolated write and OpenGL rendering methods in JPEG image sequence decoding and rendering, the problem of serialization of decoding and rendering in the prior art is solved, and efficient parallel processing and image rendering are achieved.

WO2025124160A1PCT designated stage expired Publication Date: 2025-06-19CHINA TELECOM CLOUD TECH CO LTD

Patent Information

Application Number
PCT/CN2024/135490
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-12
Filing Date
2024-11-29
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

In the prior art, the decoding and rendering process of JPEG image sequences is serialized, and CPU and GPU resources cannot be fully utilized, resulting in inefficiency.

Method used

CPU concurrent programming and thread allocation strategy are used for Huffman decoding, combined with isolated write and video memory allocation strategy, parallel processing and data writing to GPU video memory are performed, and image sequence rendering is used using OpenGL rendering method.

Benefits of technology

The decoding process of JPEG image sequences is significantly accelerated, overall performance is improved, processing time is reduced, and efficient image rendering is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024135490_19062025_PF_FP_ABST
    Figure CN2024135490_19062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of cloud computing, and specifically relates to a method for parallel decoding and rendering of a sequence of JPEG images. The method comprises the following step: on the basis of CPU concurrent programming, using a thread allocation policy to perform Huffman decoding on received JPEG images, so as to generate a sequence of frames processed in parallel. In the present application, by using CPU concurrent programming and a thread allocation policy, the efficiency of Huffman decoding is greatly improved and the time taken for preliminary processing is reduced; by means of independent decoding of threads, the throughput and efficiency are significantly improved; furthermore, in view of an isolation writing policy and a video memory allocation policy, data is efficiently written into a GPU video memory, thereby reducing data operations and increasing the speed; the parallel computing capability of a GPU makes IDCT and YUV-to-TGB conversion rapid and efficient, thereby reducing the delay and improving the output quality; by means of a shared video memory buffer policy, the utilization of hardware resources is optimized; and by using OpenGL rendering, the smoothness of rendering and the visual effect are ensured. On the basis of a combination of these optimizations, the speed and quality of image processing are significantly improved, thereby laying a solid foundation for real-time applications.
Need to check novelty before this filing date? Find Prior Art

Description

A parallel decoding and rendering method for JPEG image sequences

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on December 12, 2023, with application number 202311703765.7 and invention name “A method for parallel decoding and rendering of JPEG image sequences”, the entire contents of which are incorporated by reference into this application. Technical Field

[0003] The present application relates to the field of cloud computing technology, and in particular to a method for parallel decoding and rendering of JPEG image sequences. Background Art

[0004] Cloud desktops are a major application of cloud computing. Cloud servers send desktop content to clients in the form of image or video streams. After decoding and rendering, the client can view the desktop content, providing a user experience similar to that of a local desktop. JPEG is a lossy image encoding format with a high compression ratio. Using JPEG to transmit image data can save bandwidth and improve bitrates to a certain extent.

[0005] However, decoding and rendering JPEG image sequences on the client side is a relatively time-consuming operation. For one thing, JPEG uses variable-length Huffman coding, which makes parallel decoding difficult, requiring serial decoding. Furthermore, JPEG uses the DCT transform to change the layout of high- and low-frequency signals in image blocks. Its inverse transform is a complex convolution operation that requires significant computational time. In addition to decoding, the image must be displayed, which involves copying data from main memory to GPU memory, often resulting in significant performance overhead.

[0006] When decoding a JPEG image sequence, the typical approach is to load the images sequentially into a queue, decode them on the CPU, upload the resulting RGB data to video memory, and render them using a graphics API (D3D, OpenGL, etc.). This is a serial approach that fails to fully utilize hardware resources such as the CPU and GPU. Typical optimization and acceleration methods are targeted at JPEG image decoding, such as using assembly language for acceleration, utilizing GPUs for parallel IDCT transforms, or adding delimiters to Huffman codes during encoding to enable parallel decoding. These methods have certain limitations. Firstly, they mostly optimize for a single image and fail to consider image sequences. Secondly, they fail to optimize the rendering process. Summary of the Invention

[0007] The purpose of this application is to address the shortcomings of the prior art and to propose a parallel decoding and rendering method for JPEG image sequences.

[0008] To achieve the above objectives, the present application adopts the following technical solution: a method for parallel decoding and rendering of JPEG image sequences, comprising the following steps:

[0009] S1: Based on CPU concurrent programming, a thread allocation strategy is adopted to perform Huffman decoding on the received JPEG image and generate a parallel processing frame sequence;

[0010] S2: Based on the parallel processing frame sequence, an isolated writing strategy is adopted to write the decoded data into the GPU memory to generate a decoded frame sequence;

[0011] S3: For GPU memory management, adopt memory allocation strategy, perform memory pre-allocation, and memory mapping data sequence;

[0012] S4: Based on the video memory mapping data sequence, a macroblock segmentation algorithm is used to perform parallel processing to generate a macroblock data sequence;

[0013] S5: Based on the macroblock data sequence, using the IDCT transformation algorithm, performing pixel matrix restoration to generate two-dimensional pixel matrix image data;

[0014] S6: Based on the two-dimensional pixel matrix image data, a YUV to RGB algorithm is used to convert pixel values ​​to generate an RGB pixel value data sequence;

[0015] S7: Based on the RGB pixel value data sequence, an OpenGL rendering method is used to perform sequence rendering to generate a final rendered image sequence.

[0016] As an optional solution of the present application, the thread allocation strategy specifically refers to pre-allocating a thread for each JPEG image to be decoded, the isolated write strategy specifically refers to maintaining the isolation of each frame of image data during the writing process, the video memory allocation strategy includes pre-allocating a header and a body, the header is used to store auxiliary decoding data including the image sequence number, width, height, and quantization parameter table, and the body is used to store the decoded image data. The pixel matrix recovery specifically refers to restoring the one-dimensional image data into a two-dimensional pixel matrix form, the pixel value conversion specifically refers to converting the YUV pixel value of each macroblock into an RGB pixel value, and the sequence rendering specifically refers to using the OpenGLAPI to render the RGB data in the buffer in the order of the image sequence.

[0017] As an optional solution of the present application, the parallel processing frame sequence specifically refers to that multiple frames of JPEG images are decoded and processed in parallel at the same time, the decoded frame sequence specifically stores the Huffman decoded data in the GPU video memory, and the macroblock data sequence specifically divides the image data into multiple groups of 8×8 macroblocks.

[0018] As an optional solution of this application, based on CPU concurrent programming, a thread allocation strategy is adopted to perform Huffman decoding on the received JPEG image and generate a parallel processing frame sequence. The specific steps are as follows:

[0019] S101: Based on the CPU concurrent programming model, a thread pool scheduling algorithm is used to generate a decoding task sequence;

[0020] S102: Based on the decoding task sequence, a task priority sorting algorithm is used to generate a prioritized decoding task sequence;

[0021] S103: Based on the priority-sorted decoding task sequence, a Huffman decoding algorithm is used to generate a Huffman decoding frame sequence;

[0022] S104: Generate a parallel processing frame sequence based on the Huffman decoding frame sequence and using a parallel execution framework.

[0023] As an optional solution of the present application, based on the parallel processing of the frame sequence, an isolated writing strategy is adopted to write the decoded data into the GPU memory, and the steps of generating the decoded frame sequence are specifically as follows:

[0024] S201: Based on the parallel processing frame sequence, using memory isolation technology, generate an isolated video memory write sequence;

[0025] S202: Based on the isolated video memory write sequence, a data stream synchronization algorithm is used to generate a synchronized video memory data sequence;

[0026] S203: Based on the synchronized video memory data sequence, adopt a video memory optimization strategy to generate an optimized video memory layout sequence;

[0027] S204: Based on the optimized video memory layout sequence, a GPU parallel writing technology is used to generate a decoded frame sequence.

[0028] As an optional solution of this application, for GPU video memory management, a video memory allocation strategy is adopted to perform video memory pre-allocation. The steps of video memory mapping data sequence are specifically as follows:

[0029] S301: Based on GPU resource requirements, a dynamic memory allocation algorithm is used, and through resource requirement analysis, memory resources are dynamically allocated to generate a dynamic memory allocation sequence;

[0030] S302: Based on the dynamic video memory allocation sequence, using video memory pre-allocation technology and video memory resource mapping, perform video memory pre-allocation to generate a pre-allocated video memory address mapping sequence;

[0031] S303: Based on the pre-allocated video memory address mapping sequence, a video memory defragmentation algorithm is used, and a video memory reorganization strategy is used to defragment the video memory to generate an optimized video memory mapping sequence;

[0032] S304: Based on the optimized video memory mapping sequence, a video memory compression algorithm is adopted and a data compression technology is used to compress the video memory data to generate a video memory mapping data sequence.

[0033] As an optional solution of the present application, based on the memory mapping data sequence, a macroblock segmentation algorithm is used to perform parallel processing to generate a macroblock data sequence. Specifically, the steps are as follows:

[0034] S401: Based on the video memory mapping data sequence, an image preprocessing algorithm is used and a data decompression technology is used to generate a decompressed image data sequence;

[0035] S402: Based on the decompressed image data sequence, using a macroblock segmentation algorithm and a fixed-size segmentation technique, the image data is divided into macroblocks to generate an initial macroblock data sequence;

[0036] S403: Based on the initial macroblock data sequence, adopting a GPU parallel processing strategy and using a multi-threaded processing technology to generate a macroblock sequence after parallel processing;

[0037] S404: Based on the macroblock sequence after parallel processing, a boundary optimization algorithm is adopted and an overlapping area analysis technology is used to optimize the macroblock boundaries to generate a macroblock data sequence.

[0038] As an optional solution of the present application, based on the macroblock data sequence, the IDCT transform algorithm is used to perform pixel matrix restoration to generate two-dimensional pixel matrix image data. Specifically, the steps are as follows:

[0039] S501: Based on the macroblock data sequence, using an inverse discrete cosine transform algorithm to restore a pixel matrix and generate an IDCT transformed pixel matrix sequence;

[0040] S502: Based on the IDCT transformed pixel matrix sequence, a sharpening algorithm and a noise reduction algorithm are used to perform image detail enhancement and noise reduction processing to generate a processed pixel matrix sequence;

[0041] S503: performing data reconstruction and color space conversion based on the processed pixel matrix sequence, and generating corrected pixel matrix image data;

[0042] S504: Based on the corrected pixel matrix image data, perform data reconstruction, prepare for color conversion, and generate complete two-dimensional pixel matrix image data.

[0043] As an optional solution of the present application, based on the two-dimensional pixel matrix image data, a YUV to RGB algorithm is used to convert pixel values ​​to generate an RGB pixel value data sequence. Specifically, the steps are as follows:

[0044] S601: Based on the two-dimensional pixel matrix image data, convert YUV to RGB using a color space conversion algorithm, and generate a preliminary RGB pixel value data sequence;

[0045] S602: performing color calibration based on the preliminary RGB pixel value data sequence, and generating a calibrated RGB pixel value data sequence;

[0046] S603: Performing pixel synchronization processing based on the calibrated RGB pixel value data sequence to prevent image tearing, and generating a synchronized RGB pixel value data sequence;

[0047] S604: Based on the synchronized RGB pixel value data sequence, data packaging and buffering are performed to prepare for image rendering, and a final RGB pixel value data sequence is generated.

[0048] As an optional solution of the present application, based on the RGB pixel value data sequence, the OpenGL rendering method is used to perform sequence rendering to generate the final rendered image sequence. Specifically, the steps are as follows:

[0049] S701: Based on the RGB pixel value data sequence, using data caching technology, storing it in graphics hardware and generating buffered RGB data;

[0050] S702: Based on the buffered RGB data, use a vertex shading algorithm to perform preliminary rendering of the image and generate preliminary rendered image data;

[0051] S703: Based on the initial rendered image data, apply a fragment shading technique to complete detail rendering of the image and generate fragment rendered image data;

[0052] S704: Based on the fragment rendering image data, perform an OpenGL rendering process, output an image, and obtain a final rendered image sequence.

[0053] Compared with the prior art, the advantages and positive effects of this application are:

[0054] In this application, CPU concurrent programming and thread allocation strategies effectively accelerate Huffman decoding, significantly saving initial processing time. This acceleration provides a key improvement to overall performance. Decoding is parallelized on the CPU, and each thread decodes independently, greatly enhancing throughput and efficiency. Combined with isolated writes and video memory allocation strategies, data is efficiently and smoothly written to the GPU video memory. This strategy optimizes video memory utilization, reduces redundant data operations, and further increases speed. The GPU can simultaneously perform large-scale IDCT transformations and YUV to RGB conversions. Because the GPU is designed for parallel computing, it is far more efficient than the CPU when performing such intensive calculations. This parallel processing means quickly processing large amounts of data, reducing latency, and maintaining high-quality output. The decoded data is rendered directly on the GPU, and the shared memory buffer strategy reduces redundant copies and maximizes hardware resource utilization. Through OpenGL rendering, efficient data preparation ensures smooth, high-quality rendering. These optimizations have greatly improved the quality and speed of the image processing process, providing technical support for real-time image processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 is a schematic diagram of the main steps of this application;

[0056] Figure 2 is a schematic diagram of the refinement of S1 of this application;

[0057] FIG3 is a schematic diagram of the refinement of S2 of the present application;

[0058] FIG4 is a schematic diagram of the refinement of S3 of this application;

[0059] FIG5 is a schematic diagram of the refinement of S4 of the present application;

[0060] FIG6 is a schematic diagram of the refinement of S5 of the present application;

[0061] FIG7 is a schematic diagram of the refinement of S6 of the present application;

[0062] FIG8 is a schematic diagram of the refinement of S7 of the present application. DETAILED DESCRIPTION

[0063] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0064] In the description of this application, it should be understood that the terms "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, they should not be understood as limiting this application. In addition, in the description of this application, "multiple" means two or more, unless otherwise clearly and specifically defined.

[0065] Example 1

[0066] Referring to FIG1 , the present application provides a technical solution: a method for parallel decoding and rendering of a JPEG image sequence, comprising the following steps:

[0067] S1: Based on CPU concurrent programming, a thread allocation strategy is adopted to perform Huffman decoding on the received JPEG image and generate a parallel processing frame sequence;

[0068] S2: Based on parallel processing of frame sequences, an isolated write strategy is adopted to write decoded data into the GPU memory to generate a decoded frame sequence;

[0069] S3: For GPU memory management, adopt memory allocation strategy, perform memory pre-allocation, and memory mapping data sequence;

[0070] S4: Based on the video memory mapping data sequence, a macroblock segmentation algorithm is used to perform parallel processing to generate a macroblock data sequence;

[0071] S5: Based on the macroblock data sequence, the IDCT transform algorithm is used to restore the pixel matrix and generate two-dimensional pixel matrix image data;

[0072] S6: Based on the two-dimensional pixel matrix image data, a YUV to RGB algorithm is used to convert pixel values ​​to generate an RGB pixel value data sequence;

[0073] S7: Based on the RGB pixel value data sequence, an OpenGL rendering method is used to perform sequence rendering to generate a final rendered image sequence.

[0074] The thread allocation strategy specifically refers to pre-allocating a thread for each JPEG image to be decoded. The isolated write strategy specifically maintains the isolation of each frame of image data during the writing process. The video memory allocation strategy includes pre-allocating the header and body. The header is used to store auxiliary decoding data including the image sequence number, width, height, and quantization parameter table. The body is used to store the decoded image data. The pixel matrix recovery specifically restores the one-dimensional image data into a two-dimensional pixel matrix form. The pixel value conversion specifically converts the YUV pixel value of each macroblock into RGB pixel value. The sequence rendering specifically uses the OpenGLAPI to render the RGB data in the buffer in the order of the image sequence.

[0075] Parallel processing of frame sequences specifically refers to decoding and processing multiple JPEG images in parallel at the same time. The decoded frame sequence specifically refers to storing the Huffman decoded data in the GPU memory. The macroblock data sequence specifically refers to dividing the image data into multiple groups of 8×8 macroblocks.

[0076] First, parallel Huffman decoding of JPEG images is performed through CPU concurrent programming and thread allocation, significantly accelerating processing speed. This approach fully utilizes the computing power of multi-core CPUs, enabling each core to work simultaneously, significantly improving decoding efficiency. The strategy of assigning each thread to a single JPEG image ensures data independence during the decoding process, avoiding data contention between threads, reducing thread synchronization overhead, and further improving system responsiveness.

[0077] Secondly, efficient data management is achieved through isolated write and memory allocation strategies for GPU memory utilization. Isolated writes ensure the independence of each frame's image data during parallel processing, avoiding write conflicts. Pre-allocation of memory, by dividing it into header and body memory, not only optimizes memory usage but also reduces the overhead and fragmentation associated with dynamic memory allocation, which is particularly important for real-time image processing applications.

[0078] The image is processed in parallel using macroblock segmentation and IDCT transform algorithms, and the pixel matrix is ​​restored. Furthermore, pixel values ​​are converted using a YUV-to-RGB algorithm. The precise implementation of these algorithms ensures high fidelity from encoding to decoding. The restored two-dimensional pixel matrix image data maintains the original color and detail of the JPEG image, while conversion to RGB format facilitates direct rendering on display devices.

[0079] Finally, the OpenGL rendering method not only provides hardware-accelerated image rendering capabilities, but also takes advantage of the efficiency of modern graphics APIs, making the final image sequence rendering smooth and high-quality, which means users can enjoy richer and more realistic visual effects.

[0080] Referring to Figure 2, based on CPU concurrent programming and adopting a thread allocation strategy, the steps for performing Huffman decoding on a received JPEG image and generating a parallel processing frame sequence are as follows:

[0081] S101: Based on the CPU concurrent programming model, a thread pool scheduling algorithm is used to generate a decoding task sequence;

[0082] S102: Based on the decoding task sequence, a task priority sorting algorithm is used to generate a prioritized decoding task sequence;

[0083] S103: Based on the priority-sorted decoding task sequence, a Huffman decoding algorithm is used to generate a Huffman decoding frame sequence;

[0084] S104: Generate a parallel processing frame sequence based on the Huffman decoding frame sequence and using a parallel execution framework.

[0085] First, by employing a thread pool scheduling algorithm to generate a sequence of decoding tasks, thread resources can be effectively managed, improving system performance and responsiveness. Second, by employing a task prioritization algorithm to sort the decoding task sequence, important tasks are prioritized, improving overall decoding efficiency and user experience. Finally, the Huffman decoding algorithm is employed to decode each task, reducing image data storage space and transmission bandwidth, thereby improving image compression. Finally, a parallel execution framework is employed to simultaneously assign multiple decoding tasks to different threads for processing, further improving decoding efficiency and speed, leveraging the advantages of multi-core CPUs to accelerate the entire decoding process.

[0086] Please refer to FIG3 . Based on the parallel processing of the frame sequence, the isolated writing strategy is adopted to write the decoded data into the GPU memory. The specific steps of generating the decoded frame sequence are as follows:

[0087] S201: Based on the parallel processing of the frame sequence, an isolated video memory write sequence is generated using memory isolation technology;

[0088] S202: Based on the isolated video memory write sequence, a data stream synchronization algorithm is used to generate a synchronized video memory data sequence;

[0089] S203: Based on the synchronized video memory data sequence, adopt a video memory optimization strategy to generate an optimized video memory layout sequence;

[0090] S204: Based on the optimized video memory layout sequence, a GPU parallel writing technology is used to generate a decoded frame sequence.

[0091] First, memory isolation technology is used to generate isolated video memory write sequences, avoiding data conflicts and race conditions between different frames and improving data consistency and reliability. Second, a data stream synchronization algorithm is used to generate synchronized video memory data sequences, ensuring that the data of each frame is arranged in the correct order, guaranteeing the correctness and integrity of the decoded frame sequence. Then, a video memory optimization strategy is used to generate an optimized video memory layout sequence, improving video memory utilization and access efficiency, and accelerating the generation of the decoded frame sequence. Finally, GPU parallel write technology is used to generate the decoded frame sequence, leveraging the parallel computing power of the GPU to achieve efficient data transmission and processing, significantly shortening the generation time of the decoded frame sequence.

[0092] Please refer to FIG4 . For GPU memory management, a memory allocation strategy is adopted to perform memory pre-allocation. The steps of the memory mapping data sequence are as follows:

[0093] S301: Based on GPU resource requirements, a dynamic memory allocation algorithm is used, and through resource requirement analysis, memory resources are dynamically allocated to generate a dynamic memory allocation sequence;

[0094] S302: Based on the dynamic video memory allocation sequence, using video memory pre-allocation technology and video memory resource mapping, perform video memory pre-allocation to generate a pre-allocated video memory address mapping sequence;

[0095] S303: Based on the pre-allocated video memory address mapping sequence, a video memory defragmentation algorithm is used, and a video memory reorganization strategy is used to defragment the video memory to generate an optimized video memory mapping sequence;

[0096] S304: Based on the optimized video memory mapping sequence, a video memory compression algorithm is adopted and a data compression technology is used to compress the video memory data to generate a video memory mapping data sequence.

[0097] First, based on GPU resource requirements, a resource demand analysis is performed to determine the required memory size and quantity for each task. Second, memory resource mapping is used to map pre-allocated memory addresses to the corresponding tasks. This avoids frequent memory allocation and release operations during task execution, improving performance and efficiency. Next, a memory reorganization strategy is implemented to consolidate and organize scattered memory fragments to reduce their impact on memory utilization. This improves memory utilization and reduces waste. Finally, data compression technology is used to reduce the size of memory data, saving memory space and improving data transmission efficiency.

[0098] Referring to FIG5 , based on the memory mapping data sequence, the macroblock segmentation algorithm is used to perform parallel processing to generate the macroblock data sequence. Specifically, the steps are as follows:

[0099] S401: Based on the video memory mapping data sequence, an image preprocessing algorithm is used, and a data decompression technology is used to generate a decompressed image data sequence;

[0100] S402: Based on the decompressed image data sequence, the image data is divided into macroblocks using a macroblock segmentation algorithm and a fixed-size segmentation technique to generate an initial macroblock data sequence;

[0101] S403: Based on the initial macroblock data sequence, a GPU parallel processing strategy is adopted and a multi-threaded processing technology is used to generate a macroblock sequence after parallel processing;

[0102] S404: Based on the macroblock sequence after parallel processing, a boundary optimization algorithm is adopted and an overlapping area analysis technology is used to optimize the macroblock boundary to generate a macroblock data sequence.

[0103] First, data decompression technology is used to decompress the compressed image data in the video memory into the original image data sequence. Next, fixed-size partitioning technology is used to divide the image data into multiple macroblocks of equal size. Each macroblock contains a certain number of pixel values. Then, multi-threaded processing technology is used to assign each macroblock to a different thread for processing. Each thread is responsible for processing the data of a macroblock. During parallel processing, boundary optimization algorithms can be used to optimize macroblock boundaries. Overlapping area analysis technology is used to determine the overlapping areas between adjacent macroblocks and perform corresponding optimization processing. This can improve processing efficiency and accuracy. Finally, based on the macroblock sequence after parallel processing, a final macroblock data sequence is generated.

[0104] Referring to FIG6 , the steps of performing pixel matrix restoration and generating two-dimensional pixel matrix image data using the IDCT transform algorithm based on the macroblock data sequence are as follows:

[0105] S501: Based on the macroblock data sequence, an inverse discrete cosine transform algorithm is used to restore the pixel matrix and generate an IDCT transformed pixel matrix sequence;

[0106] S502: Based on the IDCT transformation of the pixel matrix sequence, a sharpening algorithm and a noise reduction algorithm are used to perform image detail enhancement and noise reduction processing to generate a processed pixel matrix sequence;

[0107] S503: Based on the processed pixel matrix sequence, perform data reconstruction and color space conversion, and generate corrected pixel matrix image data;

[0108] S504: Based on the corrected pixel matrix image data, perform data reconstruction, prepare for color conversion, and generate complete two-dimensional pixel matrix image data.

[0109] First, the IDCT transform converts the macroblock data represented in the frequency domain back into a pixel matrix sequence represented in the spatial domain. Next, a sharpening algorithm enhances image edges and detail, while a noise reduction algorithm reduces the noise level in the image. This improves image quality and clarity. Then, data reconstruction converts the pixel matrix into a complete two-dimensional array. A color space conversion transforms the image from the original color space to the target color space, ensuring that the resulting pixel matrix image data is consistent with the original image. Finally, data reconstruction converts the pixel matrix into a data format suitable for display and storage. A preparatory color conversion transforms the image from the source color space to the target color space, ensuring that the resulting complete two-dimensional pixel matrix image data can be correctly displayed and processed.

[0110] Referring to FIG. 7 , based on the two-dimensional pixel matrix image data, the steps of converting pixel values ​​using the YUV to RGB algorithm to generate an RGB pixel value data sequence are as follows:

[0111] S601: Based on the two-dimensional pixel matrix image data, a color space conversion algorithm is used to convert YUV to RGB, and a preliminary RGB pixel value data sequence is generated;

[0112] S602: performing color calibration based on the preliminary RGB pixel value data sequence, and generating a calibrated RGB pixel value data sequence;

[0113] S603: Perform pixel synchronization processing based on the calibrated RGB pixel value data sequence to prevent image tearing, and generate a synchronized RGB pixel value data sequence;

[0114] S604: Based on the synchronized RGB pixel value data sequence, data packaging and buffering are performed to prepare for image rendering and generate a final RGB pixel value data sequence.

[0115] First, converting the YUV color space to the RGB color space better accommodates the human eye's color perception. Compared to the combination of the luminance signal Y and the chrominance signals U and V, the RGB color space better aligns with our intuitive understanding of the three primary colors of red, green, and blue, and can more accurately reproduce the image's color information. Second, color calibration can further optimize image color performance. Because the color characteristics of different devices or displays may vary, calibration maps the image's colors to a standard color space, ensuring more consistent and accurate color display across devices. Furthermore, pixel synchronization can prevent image tearing. On high-refresh-rate displays, if the image rendering speed cannot keep up with the display's refresh rate, tearing can occur, where the same image frame is split into two halves and displayed separately on the screen. Pixel synchronization ensures smooth image display on the display, avoiding tearing. Finally, data packing and buffering improve image rendering efficiency. Packing and buffering pixel value data sequences can reduce data transmission times and latency, improving image rendering speed and performance. This is particularly important for applications with high real-time requirements, such as video playback and gaming.

[0116] Referring to FIG8 , the steps for performing sequence rendering based on the RGB pixel value data sequence and using the OpenGL rendering method to generate the final rendered image sequence are as follows:

[0117] S701: Based on the RGB pixel value data sequence, a data caching technology is used to store the data in the graphics hardware and generate buffered RGB data;

[0118] S702: Based on the buffered RGB data, use the vertex shading algorithm to perform preliminary rendering of the image and generate preliminary rendered image data;

[0119] S703: Based on the initial rendered image data, apply the fragment shading technology to complete the image detail rendering and generate the fragment rendered image data;

[0120] S704: Based on the fragment rendering image data, perform an OpenGL rendering process, output the image, and obtain a final rendered image sequence.

[0121] First, by using data caching technology to store RGB pixel values ​​in the graphics hardware, data reading and transmission speeds can be accelerated, data transmission delays can be reduced, and rendering efficiency can be improved. Secondly, using the vertex shading algorithm for preliminary image rendering, the vertex shader can be used to preprocess and transform the image, reducing the complexity and computational complexity of image processing and improving rendering efficiency and quality. Then, fragment shading technology is applied to complete the detailed rendering of the image, using the fragment shader to accurately calculate and draw each pixel, achieving more detailed and realistic image effects, improving the quality and realism of the rendering. Finally, the image is output through the OpenGL rendering process to obtain the final rendered image sequence. OpenGL provides a complete rendering pipeline and rich rendering functions, which can realize a variety of complex rendering effects and operations.

[0122] The above are merely preferred embodiments of the present application and do not constitute other forms of limitation to the present application. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification of the above embodiments made according to the technical essence of the present application without departing from the content of the technical solution of the present application shall still fall within the scope of protection of the technical solution of the present application.

Claims

1. A method for parallel decoding and rendering of a JPEG image sequence, characterized in that: The following steps are involved: Based on CPU concurrent programming, thread allocation strategy is adopted to perform Huffman decoding on the received JPEG images and generate parallel processing frame sequences; Based on the parallel processing frame sequence, an isolated writing strategy is adopted to write the decoded data into the GPU memory to generate a decoded frame sequence; For GPU video memory management, a video memory allocation strategy is adopted to perform video memory pre-allocation and video memory mapping data sequence; Based on the video memory mapping data sequence, a macroblock segmentation algorithm is used to perform parallel processing to generate a macroblock data sequence; Based on the macroblock data sequence, an IDCT transform algorithm is used to perform pixel matrix restoration to generate two-dimensional pixel matrix image data; Based on the two-dimensional pixel matrix image data, a YUV to RGB algorithm is used to perform pixel value conversion to generate an RGB pixel value data sequence; Based on the RGB pixel value data sequence, an OpenGL rendering method is used to perform sequence rendering to generate a final rendered image sequence.

2. The method for parallel decoding and rendering of JPEG image sequences according to claim 1, characterized in that: The thread allocation strategy specifically refers to pre-allocating a thread for each JPEG image to be decoded, the isolated write strategy specifically refers to maintaining the isolation of each frame of image data during the writing process, the video memory allocation strategy includes pre-allocating a header and a body, the header is used to store auxiliary decoding data including an image sequence number, width, height, and a quantization parameter table, and the body is used to store decoded image data, the pixel matrix recovery specifically refers to restoring one-dimensional image data into a two-dimensional pixel matrix form, the pixel value conversion specifically refers to converting the YUV pixel value of each macroblock into an RGB pixel value, and the sequence rendering specifically refers to using the OpenGL API to render the RGB data in the buffer in the order of the image sequence.

3. The method for parallel decoding and rendering of JPEG image sequences according to claim 1, characterized in that: The parallel processing frame sequence specifically refers to that multiple frames of JPEG images are decoded and processed in parallel at the same time, the decoded frame sequence specifically stores the Huffman decoded data in the GPU video memory, and the macroblock data sequence specifically divides the image data into multiple groups of 8×8 macroblocks.

4. The method for parallel decoding and rendering of JPEG image sequences according to claim 1, characterized in that: Based on CPU concurrent programming, thread allocation strategy is adopted to perform Huffman decoding on the received JPEG image and generate parallel processing frame sequence. The specific steps are as follows: Based on the CPU concurrent programming model, the thread pool scheduling algorithm is used to generate the decoding task sequence; Based on the decoding task sequence, a task priority sorting algorithm is used to generate a decoding task sequence after priority sorting; Based on the priority-sorted decoding task sequence, a Huffman decoding algorithm is used to generate a Huffman decoding frame sequence; Based on the Huffman decoded frame sequence and using a parallel execution framework, a parallel processing frame sequence is generated.

5. The method for parallel decoding and rendering of JPEG image sequences according to claim 1, characterized in that: Based on the parallel processing frame sequence, an isolated writing strategy is adopted to write the decoded data into the GPU memory, and the steps of generating the decoded frame sequence are specifically as follows: Based on the parallel processing frame sequence, a memory isolation technology is used to generate an isolated video memory write sequence; Based on the isolated video memory write sequence, a data stream synchronization algorithm is used to generate a synchronized video memory data sequence; Based on the synchronized video memory data sequence, a video memory optimization strategy is adopted to generate an optimized video memory layout sequence; Based on the optimized video memory layout sequence, a GPU parallel writing technology is adopted to generate a decoded frame sequence.

6. The method for parallel decoding and rendering of JPEG image sequences according to claim 1, characterized in that: For the management of GPU video memory, a video memory allocation strategy is adopted to pre-allocate video memory. The steps of video memory mapping data sequence are as follows: Based on GPU resource requirements, a dynamic video memory allocation algorithm is used. Through resource requirement analysis, video memory resources are dynamically allocated to generate a dynamic video memory allocation sequence. Based on the dynamic video memory allocation sequence, a video memory pre-allocation technology is adopted, and video memory resource mapping is used to perform video memory pre-allocation to generate a pre-allocated video memory address mapping sequence; Based on the pre-allocated video memory address mapping sequence, a video memory defragmentation algorithm is adopted, and a video memory reorganization strategy is used to defragment the video memory to generate an optimized video memory mapping sequence; Based on the optimized video memory mapping sequence, a video memory compression algorithm is adopted, and a data compression technology is used to perform compression processing on the video memory data to generate a video memory mapping data sequence.

7. The method for parallel decoding and rendering of JPEG image sequences according to claim 1, characterized in that: Based on the video memory mapping data sequence, a macroblock segmentation algorithm is used to perform parallel processing to generate a macroblock data sequence in the following specific steps: Based on the video memory mapping data sequence, an image preprocessing algorithm is adopted, and a decompressed image data sequence is generated through a data decompression technology; Based on the decompressed image data sequence, a macroblock segmentation algorithm is used, and a fixed-size partitioning technique is used to perform macroblock partitioning of the image data to generate an initial macroblock data sequence; Based on the initial macroblock data sequence, a GPU parallel processing strategy is adopted, and a parallel processed macroblock sequence is generated through multi-thread processing technology; Based on the macroblock sequence after parallel processing, a boundary optimization algorithm is adopted, and the overlapping area analysis technology is used to optimize the macroblock boundary to generate a macroblock data sequence.

8. The method for parallel decoding and rendering of JPEG image sequences according to claim 1, characterized in that: Based on the macroblock data sequence, the IDCT transform algorithm is used to restore the pixel matrix and generate the two-dimensional pixel matrix image data in the following steps: Based on the macroblock data sequence, an inverse discrete cosine transform algorithm is used to restore the pixel matrix and generate an IDCT transformed pixel matrix sequence; Based on the IDCT transformed pixel matrix sequence, a sharpening algorithm and a noise reduction algorithm are used to perform image detail enhancement and noise reduction processing to generate a processed pixel matrix sequence; Based on the processed pixel matrix sequence, data reconstruction and color space conversion are performed to generate corrected pixel matrix image data; Based on the corrected pixel matrix image data, data reconstruction is performed, color conversion is prepared, and complete two-dimensional pixel matrix image data is generated.

9. The method for parallel decoding and rendering of JPEG image sequences according to claim 1, characterized in that: Based on the two-dimensional pixel matrix image data, the steps of converting pixel values ​​using a YUV to RGB algorithm to generate an RGB pixel value data sequence are specifically as follows: Based on the two-dimensional pixel matrix image data, a color space conversion algorithm is used to convert YUV to RGB, and a preliminary RGB pixel value data sequence is generated; Based on the preliminary RGB pixel value data sequence, color calibration is performed to generate a calibrated RGB pixel value data sequence; Based on the calibrated RGB pixel value data sequence, pixel synchronization processing is performed to prevent image tearing, and a synchronized RGB pixel value data sequence is generated; Based on the synchronized RGB pixel value data sequence, data packaging and buffering are performed to prepare for image rendering, and a final RGB pixel value data sequence is generated.

10. The method for parallel decoding and rendering of JPEG image sequences according to claim 1, characterized in that: Based on the RGB pixel value data sequence, the OpenGL rendering method is used to perform sequence rendering to generate the final rendered image sequence. The specific steps are: Based on the RGB pixel value data sequence, a data caching technique is used to store the data in the graphics hardware and generate buffered RGB data; Based on the buffered RGB data, using a vertex shading algorithm, perform preliminary rendering of the image and generate preliminary rendering image data; Based on the initial rendered image data, applying a fragment shading technique to complete image detail rendering and generate fragment rendered image data; Based on the fragment rendering image data, an OpenGL rendering process is performed, an image is output, and a final rendered image sequence is obtained.

Citation Information

Patent Citations

  • JPEG high-speed decoding method

    CN106559674A

  • High efficiency decoding and playing method and high efficiency decoding and playing system used for high definition videos

    CN108366288A

  • Video frame adaptive rendering processing method, computer device and storage medium

    CN116912385A

  • JPEG image sequence parallel decoding and rendering method

    CN117834903A

  • Keyframe-based video codec designed for GPU decoding

    US20180199067A1

Cited By

  • Adaptive frame abstraction interface management method and system for multiple coding formats

    CN122395338A