Image decoding method, encoding method, electronic equipment and readable storage medium

By decoupling the image decoding process and adopting asynchronous independent clock processing, the problem of limited image encoding and decoding efficiency and throughput in the prior art is solved, efficient image decoding is achieved and device memory utilization is improved.

CN120091142APending Publication Date: 2025-06-03SHANGHAI YUNSUI TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510246881.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

In the existing image encoding and decoding technology, the entropy decoding and codec bit context dependence makes the hardware timing difficult to converge and the working frequency cannot be improved, resulting in limited image encoding and decoding efficiency and throughput.

Method used

By decoupling the image decoding process, the pixel residual information is obtained by using the first decoding sub-stage processing, and buffering it into a shared macroblock information buffer, the pixel residual information to be processed is obtained by using an asynchronous independent clock to perform the second decoding sub-stage processing, including inverse quantization, inverse discrete cosine transformation and pixel code stream writing back.

Benefits of technology

It realizes asynchronous parallel processing of the decoding process, improves image decoding efficiency and throughput, solves the bit dependence bottleneck of code streams in the entropy decoding process, and improves the memory utilization of the device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120091142A_ABST
    Figure CN120091142A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an image decoding method, an image coding method, electronic equipment and a readable storage medium. The decoding method comprises the following steps: during image decoding, performing first decoding sub-stage processing on a to-be-decoded image code stream to obtain pixel residual information; wherein the first decoding sub-stage comprises the steps of code stream acquisition and entropy decoding of an image to be decoded; caching the pixel residual error information to a shared macro block information buffer; when an image is decoded, acquiring residual information of pixels to be processed in the macro block information buffer by adopting an asynchronous independent clock; performing second decoding sub-stage processing on the to-be-processed pixel residual information to obtain image decoding information; wherein the second decoding sub-stage comprises inverse quantization, inverse discrete cosine transform and pixel code stream write-back. By decoupling the decoding process, asynchronous parallel processing of the decoding process can be realized, the image decoding efficiency is improved, and the memory utilization rate of equipment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to an image decoding method, an encoding method, an electronic device, and a readable storage medium. Background Art

[0002] With the large-scale application of Chinese text-to-image or image-to-image large models of artificial intelligence generated content (AIGC), as well as traditional vision models, the requirements for the latency and throughput of image encoding and decoding in model processing are getting higher and higher.

[0003] Existing image encoding and decoding are at the pipeline level. During the encoding and decoding process, they are limited by entropy decoding and the context dependence of encoding and decoding bits, resulting in difficult convergence of hardware timing, inability to increase the hardware operating frequency, and serious impact on image encoding and decoding efficiency and throughput. Summary of the Invention

[0004] The present invention provides an image decoding method, an encoding method, an electronic device, and a readable storage medium to improve image decoding efficiency and throughput through the decoupling of the decoding process.

[0005] According to one aspect of the present invention, an image decoding method is provided, and the method includes:

[0006] During image decoding, perform a first decoding sub-stage process on the image bitstream to be decoded to obtain pixel residual information; wherein, the first decoding sub-stage includes: bitstream acquisition of the image to be decoded and entropy decoding;

[0007] Cache the pixel residual information into a shared macroblock information cache;

[0008] During image decoding, asynchronously and independently use a clock to obtain the pixel residual information to be processed in the macroblock information cache;

[0009] Perform a second decoding sub-stage process on the pixel residual information to be processed to obtain image decoding information; wherein, the second decoding sub-stage includes inverse quantization, inverse discrete cosine transform, and pixel bitstream write-back.

[0010] According to one aspect of the present invention, an image encoding method is provided, including:

[0011] During image encoding, perform a first encoding sub-stage process on the image bitstream to be encoded to obtain quantized residual information; wherein, the first encoding sub-stage includes: pixel acquisition of the image to be encoded, discrete cosine transform, and quantization;

[0012] Cache the quantized residual information into a shared macroblock information cache;

[0013] When performing image encoding, an asynchronous independent clock is used to obtain the quantization residual information to be processed in the macroblock information buffer;

[0014] The quantization residual information to be processed is subjected to a second encoding sub-stage process to obtain an image encoding bitstream; wherein, the second encoding sub-stage includes: entropy encoding and writing back the image encoding bitstream.

[0015] According to another aspect of the present invention, an image decoding device is provided, which includes:

[0016] A first decoding sub-stage processing module, configured to perform a first decoding sub-stage process on the bitstream of the image to be decoded during image decoding to obtain pixel residual information; wherein, the first decoding sub-stage includes: obtaining the bitstream of the image to be decoded and entropy decoding;

[0017] A pixel residual information buffer module, configured to buffer the pixel residual information into a shared macroblock information buffer;

[0018] A module for obtaining the pixel residual information to be processed, configured to use an asynchronous independent clock to obtain the pixel residual information to be processed in the macroblock information buffer during image decoding;

[0019] A second decoding sub-stage processing module, configured to perform a second decoding sub-stage process on the pixel residual information to be processed to obtain image decoding information; wherein, the second decoding sub-stage includes: inverse quantization, inverse discrete cosine transform, and writing back the pixel bitstream.

[0020] According to another aspect of the present invention, an image encoding device is provided, which includes:

[0021] A first encoding sub-stage processing module, configured to perform a first encoding sub-stage process on the bitstream of the image to be encoded during image encoding to obtain quantization residual information; wherein, the first encoding sub-stage includes: obtaining pixels of the image to be encoded, discrete cosine transform, and quantization;

[0022] A quantization residual information buffer module, configured to buffer the quantization residual information into a shared macroblock information buffer;

[0023] A module for obtaining the quantization residual information to be processed, configured to use an asynchronous independent clock to obtain the quantization residual information to be processed in the macroblock information buffer during image encoding;

[0024] A second encoding sub-stage processing module, configured to perform a second encoding sub-stage process on the quantization residual information to be processed to obtain an image encoding bitstream; wherein, the second encoding sub-stage includes: entropy encoding and writing back the image encoding bitstream.

[0025] According to another aspect of the present invention, there is provided an electronic device, the electronic device comprising:

[0026] at least one processor; and

[0027] a memory communicatively connected to the at least one processor; wherein,

[0028] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the image decoding method or the image encoding method according to any embodiment of the present invention.

[0029] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the image decoding method or the image encoding method according to any embodiment of the present invention when executed.

[0030] According to another aspect of the present invention, there is provided a computer program product comprising a computer program which, when executed by a processor, implements the image decoding method or the image encoding method according to any embodiment of the present invention.

[0031] In the technical solution of the embodiment of the present invention, during image decoding, the image bitstream to be decoded is processed in a first decoding sub-stage to obtain pixel residual information; wherein, the first decoding sub-stage includes: bitstream acquisition of the image to be decoded and entropy decoding; caching the pixel residual information into a shared macroblock information buffer; during image decoding, asynchronously independent clocks are used to obtain the pixel residual information to be processed in the macroblock information buffer; and the pixel residual information to be processed is processed in a second decoding sub-stage to obtain image decoding information; wherein, the second decoding sub-stage includes inverse quantization, inverse discrete cosine transform, and pixel bitstream write-back, which solves the problem of bit dependency bottleneck in the decoding process. By decoupling the decoding process, asynchronous parallel processing of the decoding process can be achieved, improving the image decoding efficiency and the device memory utilization rate.

[0032] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0034] Figure 1 is a flowchart of an image decoding method provided in Embodiment 1 of the present invention;

[0035] Figure 2 is a schematic diagram of the execution of an image decoding method provided in Embodiment 1 of the present invention;

[0036] Figure 3 is a schematic diagram of the register segment mapping provided in Embodiment 1 of the present invention;

[0037] Figure 4 is a flowchart of an image encoding method provided in Embodiment 2 of the present invention;

[0038] Figure 5 is a schematic diagram of the execution of an image encoding method provided in Embodiment 2 of the present invention;

[0039] Figure 6 is a flowchart of an image encoding and decoding method provided in Embodiment 3 of the present invention;

[0040] Figure 7 is a schematic diagram of the image encoding and decoding process provided in Embodiment 3 of the present invention;

[0041] Figure 8 is a schematic diagram of the structure of an image decoding device provided in Embodiment 4 of the present invention;

[0042] Figure 9 is a schematic diagram of the structure of an image encoding device provided in Embodiment 5 of the present invention;

[0043] Figure 10 is a schematic diagram of the structure of an image encoding and decoding device provided in Embodiment 6 of the present invention;

[0044] Figure 11 is a schematic diagram of the structure of an electronic device for implementing the image decoding method and / or the image encoding method of the embodiments of the present invention. Detailed implementation manners

[0045] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0046] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0047] Embodiment 1

[0048] Figure 1 is a flowchart of an image decoding method provided according to Embodiment 1 of the present invention. This embodiment is applicable to the situation of improving decoding efficiency and throughput during image decoding. This method can be executed by an image decoding device, which can be implemented in the form of hardware and / or software. The image decoding device can be configured in an electronic device, which can be a mobile phone, a tablet computer (PAD), a computer (PC, Personal Computer), etc. As Figure 1 shown, the method includes:

[0049] Step 110: During image decoding, perform a first decoding sub-stage process on the image code stream to be decoded to obtain pixel residual information.

[0050] Among them, the first decoding sub-stage includes: code stream acquisition of the image to be decoded and entropy decoding.

[0051] The process of image decoding may include: obtaining the bitstream of the image to be decoded, entropy decoding, inverse quantization, inverse discrete cosine transform, and writing back the pixel bitstream. Due to the bit dependence of the bitstream in the image decoding process, the image decoding delay is relatively high and the throughput is low. In the embodiments of the present invention, the image decoding process is decoupled and segmented to generate two decoding sub-stages. In the first decoding sub-stage, the bitstream fetching (Stream Fetch, SF) and entropy decoding (Entropy Decoding, ED) of the bitstream of the image to be decoded can be performed.

[0052] Specifically, the bitstream fetching of the image to be decoded can be to read the input bitstream of image decoding into the hardware bit first-in-first-out (Bit FIFO) in the memory such as the bitstream buffer through stream direct memory access (Stream Direct Memory Access, SDMA). Then, the bitstream can be read in the Bit FIFO for entropy decoding to obtain the pixel residual information.

[0053] To accelerate the image decoding process, when performing the processing of the first decoding sub-stage on the bitstream of the image to be decoded, the processing flow in the first decoding sub-stage can be performed in parallel.

[0054] In an alternative embodiment of the embodiments of the present invention, optionally, the processing of the first decoding sub-stage on the bitstream of the image to be decoded includes: for the bitstream of the image to be decoded with a preset number of macroblocks, when performing entropy decoding on the bitstream of the image to be decoded in the current macroblock, performing bitstream fetching on the bitstream of the image to be decoded in the next macroblock.

[0055] When performing the processing flow in the first decoding sub-stage in parallel, different data volume units can be used. For example, the processing of the first decoding sub-stage in parallel can be performed in units of macroblocks. Among them, a macroblock is a square area composed of 16×16 pixels. Figure 2 It is a schematic diagram of image decoding execution provided by Embodiment 1 of the present invention. Figure 2 MB0, MB1, MB2, MB3, and MB4 in it are different macroblocks. As Figure 2 shown, bitstream fetching and entropy decoding can be performed on the bitstream of the first image to be decoded, that is, the bitstream of the image to be decoded corresponding to MB0; when performing entropy decoding on the information corresponding to MB0, bitstream fetching can be performed on the bitstream of the second image to be decoded, that is, the bitstream of the image to be decoded corresponding to MB1; then, entropy decoding can be performed on the information corresponding to MB1, and at the same time, bitstream fetching can be performed on the bitstream of the third image to be decoded, that is, the bitstream of the image to be decoded corresponding to MB2; and so on, until the bitstream fetching of the last bitstream of the image to be decoded is completed and then its entropy decoding is performed.

[0056] By processing the first decoding sub - stage in parallel, the image decoding efficiency can be improved, achieving the image decoding effect of high density and low latency.

[0057] Step 120: Cache the pixel residual information into the shared macro - block information buffer.

[0058] After the bit - stream parsing is completed through the first decoding sub - stage, the pixel residual information can be cached into the shared macro - block information buffer. Among them, by connecting the first decoding sub - stage and the second decoding sub - stage through the shared macro - block information buffer, a complete decoding process can be realized when the decoding process is decoupled. The shared macro - block information buffer can be a buffer unit in terms of macro - blocks, which can support data access from both the first decoding sub - stage and the second decoding sub - stage simultaneously. For example, the macro - block information buffer can be a Zigzag Buffer.

[0059] Optionally, caching the pixel residual information into the shared macro - block information buffer includes: caching the pixel residual information of a preset number of macro - blocks into the shared macro - block information buffer in a zigzag scan manner with two macro - block rows as a unit.

[0060] Specifically, the Zigzag Buffer can perform a zigzag scan on the pixel residual information with two macro - block rows as a unit. As Figure 2 shown, the Zigzag Buffer completes the data interaction between the first decoding sub - stage (Pipe0) and the second decoding sub - stage (Pipe1). The Zigzag Buffer caches the intermediate macro - block information of the pipe in a zigzag scan manner and organizes the data structure of the Zigzag Buffer with two macro - block rows as a unit (i.e., 32 rows of pixels).

[0061] Through the shared macro - block information buffer, asynchronous and independent processing of the first decoding sub - stage and the second decoding sub - stage can be realized, thereby decoupling the decoding process, improving the decoding efficiency, and achieving the image decoding effect of high density and low latency.

[0062] Step 130: When decoding an image, use an asynchronous and independent clock to obtain the pixel residual information to be processed in the macro - block information buffer.

[0063] The first decoding sub - stage and the second decoding sub - stage can be executed asynchronously and independently. Therefore, when performing the second decoding sub - stage processing, an asynchronous and independent clock can be used to obtain the pixel residual information to be processed in the macro - block information buffer.

[0064] Optionally, using an asynchronous and independent clock to obtain the pixel residual information to be processed in the macro - block information buffer includes: using an asynchronous and independent clock to obtain the pixel residual information to be processed in two macro - block rows through a zigzag scan in the macro - block information buffer.

[0065] As shown Figure 2 in the figure, the pixel residual information to be processed of MB0, MB1, MB2, MB3, MB4... can be sequentially obtained through zigzag scanning in the macroblock information buffer. Specifically, in the first decoding sub-stage, the pixel residual information is cached in the shared macroblock information buffer, and in the second decoding sub-stage, the pixel residual information to be processed is obtained from the macroblock information buffer. By using asynchronous independent clocks, the image decoding process can be decoupled. The inverse quantization process during image decoding does not need to wait for the entropy decoding process to complete, thereby improving the image decoding efficiency, solving the bottleneck of bit dependence in the entropy decoding process, and greatly improving the throughput rate. Moreover, by using asynchronous independent clocks, the clock frequency of inverse quantization can be effectively increased, and the hardware area efficiency ratio of image decoding can be effectively improved.

[0066] Step 140: Perform second decoding sub-stage processing on the pixel residual information to be processed to obtain image decoding information.

[0067] Among them, the second decoding sub-stage includes inverse quantization (IQ), inverse discrete cosine transform (IDCT), and pixel write back (PWB).

[0068] IQ can obtain the pixel residual information to be processed from the shared macroblock information buffer such as Zigzag Buffer for inverse quantization. IDCT can perform inverse discrete cosine transform on the inverse quantized information to complete the pixel reconstruction function, and after pixel reconstruction, it can be output to the pixel first-in-first-out (Pixel FIFO). In PWB, the pixel value can be written back to memory such as the frame buffer. For example, the pixel value can be written back to the frame buffer through video direct memory access (VDMA).

[0069] To accelerate the image decoding process, when performing second decoding sub-stage processing on the image code stream to be decoded, the processing flow in the second decoding sub-stage can be performed in parallel.

[0070] In an optional implementation manner of the embodiment of the present invention, optionally, performing second decoding sub-stage processing on the pixel residual information to be processed includes: when performing inverse discrete cosine transform on the pixel residual information to be processed of the preset number of macroblocks in the current macroblock, performing pixel write back on the pixel residual information to be processed of the previous macroblock, and performing inverse quantization on the pixel residual information to be processed of the next macroblock.

[0071] When adopting a parallel mode for the processing flow in the second decoding sub-phase, different data volume units can be used. For example, the second decoding sub-phase in the parallel mode can be processed in units of macroblocks. As Figure 2 shown, in the second decoding sub-phase, the pixel residual information to be processed corresponding to MB0 can be dequantized, inverse discrete cosine transformed, and pixel bitstream written back; when performing inverse discrete cosine transformation on the information corresponding to MB0, the pixel residual information to be processed corresponding to MB1 can be dequantized simultaneously; when writing back the pixel bitstream for the information corresponding to MB0, the information corresponding to MB1 can be inverse discrete cosine transformed simultaneously, and the pixel residual information to be processed corresponding to MB2 can be dequantized; when writing back the pixel bitstream for the information corresponding to MB1, the information corresponding to MB2 can be inverse discrete cosine transformed simultaneously, and the pixel residual information to be processed corresponding to MB3 can be dequantized, and so on, until the dequantization and inverse discrete cosine transformation of the last pixel residual information to be processed are completed and then its pixel bitstream is written back.

[0072] By adopting a parallel mode for the processing flow in the second decoding sub-phase, the image decoding efficiency can be improved, achieving an image decoding effect with high density and low latency.

[0073] The technical solution of this embodiment, when performing image decoding, processes the image bitstream to be decoded in the first decoding sub-phase to obtain pixel residual information; wherein, the first decoding sub-phase includes: bitstream acquisition and entropy decoding of the image bitstream to be decoded; caching the pixel residual information into a shared macroblock information buffer; when performing image decoding, asynchronously and independently obtaining the pixel residual information to be processed in the macroblock information buffer using a clock; performing the second decoding sub-phase processing on the pixel residual information to be processed to obtain image decoding information; wherein, the second decoding sub-phase includes dequantization, inverse discrete cosine transformation, and pixel bitstream writing back, solving the problem of the bitstream bit dependence bottleneck during image decoding. By decoupling and splitting the image decoding pipeline and using a shared buffer to interact with the bitstream and pixel information, the bitstream bit dependence bottleneck in the entropy decoding process can be effectively solved, and the throughput can be improved; by adopting an asynchronous and independent clock domain in image decoding, the clock frequencies of the inverse discrete cosine transformation and dequantization can be effectively increased, effectively improving the hardware area efficiency ratio of image decoding.

[0074] Based on the above implementation manner, in order to reduce the address switching overhead in the image decoding scenario during data interaction and improve the memory utilization rate of the device side, optionally, the image decoding method further includes: when performing image decoding data interaction with the host, obtaining a unified virtual address and converting the unified virtual address into a device virtual address; obtaining the register segment mapping bound to the data interaction, and based on the mapping relationship between the virtual address and the physical address in the register segment mapping and the device virtual address, obtaining the device physical address; performing image decoding data interaction through the device physical address.

[0075] Among them, data interaction may refer to reading data from a memory such as a frame buffer or a bit buffer for image decoding. At this time, the address switching overhead can be reduced through register segment mapping, and the memory utilization rate of the device side can be improved.

[0076] When multiple image decoding sessions are concurrent, each image decoding can independently issue an image decoding task at the host side, and carry a unified virtual address (Unified virtual address, UVA) in the task. When the hardware decoding device receives the UVA, it can convert the UVA into a device virtual address (Device Virtual Address, DVA). Among them, DVA is an abstraction layer provided for external devices in a computer system, enabling external devices to have their own virtual address space like the CPU. A mapping relationship can be pre-established between the UVA and the DVA, and when the hardware decoding device receives the UVA, it can determine the DVA according to the mapping relationship.

[0077] In the embodiment of the present invention, in order to reduce the address switching overhead in the image decoding scenario during data interaction and improve the memory utilization rate of the device side, register segment mapping is used to replace the traditional memory management unit (Memory Management Unit, MMU) page table mapping. Among them, the mapping relationship between the virtual address and the physical address is directly constructed in the register segment mapping.

[0078] Exemplarily, Figure 3 is a schematic diagram of register segment mapping provided according to Embodiment 1 of the present invention. As Figure 3 shown, each hardware image decoding device may include 4 codec stream hardware context identifiers (Logic Session Identity, LSID). Each LSID may include 32 segments. Each segment may include a virtual base address (VA_BASE), an address size, a physical base address (PA_BASE), and an enable signal (Valid). For example, the register segment mapping may support a 64-bit virtual address space and a 48-bit physical address space. In practical applications, the virtual address space, physical address space, number of LSIDs, and number of segments in the register segment mapping can all be adjusted according to actual situations, so as to dynamically adapt to a unified address space with different address bit widths, without being restricted by the address bit width in the traditional MMU page table mapping.

[0079] According to the mapping relationship between the virtual base address and the physical base address in the register segment mapping, and the address size, the corresponding device physical address can be determined for the device virtual address. Furthermore, through the device physical address, read and write access to device memory such as frame buffer or bit buffer is performed to complete image decoding. By replacing the traditional MMU page table mapping with register segment mapping, the end-to-end latency in the synchronous decoding mode is greatly reduced, the address switching overhead in the synchronous image decoding scenario is reduced, and the device-side memory utilization rate is improved.

[0080] Embodiment 2

[0081] Figure 4 FIG. is a flowchart of an image encoding method provided according to Embodiment 2 of the present invention. This embodiment is applicable to the situation of improving the encoding efficiency and throughput during image encoding. This method can be executed by an image encoding device, which can be implemented in the form of hardware and / or software. The image encoding device can be configured in an electronic device, and the electronic device can be a mobile phone, a tablet computer (PAD), a computer (PC, Personal Computer), etc. As Figure 4 shown, the method includes:

[0082] Step 410, during image encoding, perform a first encoding sub-stage process on the image code stream to be encoded to obtain quantization residual information.

[0083] Among them, the first encoding sub-stage includes: pixel fetch (PF) of the image to be encoded, discrete cosine transform (DCT), and quantization (Q).

[0084] Image encoding is the inverse process of image decoding. The process flow of image encoding can include: pixel fetch of the image to be encoded, discrete cosine transform, quantization, entropy encoding, and writing back of the image encoding code stream. Due to the bit dependence of the code stream in the image encoding process, the image encoding latency is relatively high and the throughput is low. In the embodiment of the present invention, the image encoding process is decoupled and divided into two encoding sub-stages. In the first encoding sub-stage, pixel fetch, discrete cosine transform, and quantization of the image code stream to be encoded can be performed.

[0085] During image encoding, the pixel data of the image to be encoded can be read from the frame Buffer to the Pixel FIFO through VDMA. Then, the pixel data of the image to be encoded is read from the Pixel FIFO to the CPU processing unit for DCT processing. After that, the DCT processing result, that is, the transformed domain pixel value, is quantized.

[0086] To accelerate the image encoding process, when performing the first encoding sub-stage processing on the image bitstream to be encoded, a parallel manner can be adopted for the processing flow in the first encoding sub-stage.

[0087] In an alternative embodiment of the present invention, optionally, performing the first encoding sub-stage processing on the image bitstream to be encoded includes: for the image bitstream to be encoded with a preset number of macroblocks, when performing discrete cosine transform on the image bitstream of the current macroblock, quantizing the image bitstream of the previous macroblock, and obtaining pixels of the image bitstream of the next macroblock.

[0088] When adopting a parallel manner for the processing flow in the first encoding sub-stage, different data volume units can be used. For example, the first encoding sub-stage processing in the parallel manner can be performed in units of macroblocks. Figure 5 It is a schematic diagram of image encoding execution provided by Embodiment 2 of the present invention. As Figure 5 shown, in the first encoding sub-stage, pixel acquisition, discrete cosine transform, and quantization of the image to be encoded can be performed on the image bitstream corresponding to MB0; when performing discrete cosine transform on the information corresponding to MB0, pixel acquisition can be simultaneously performed on the image bitstream corresponding to MB1; when quantizing the information corresponding to MB0, discrete cosine transform can be simultaneously performed on the information corresponding to MB1, and pixel acquisition can be performed on the image bitstream corresponding to MB2; when quantizing the information corresponding to MB1, discrete cosine transform can be simultaneously performed on the information corresponding to MB2, and pixel acquisition can be performed on the image bitstream corresponding to MB3, and so on, until pixel acquisition and discrete cosine transform of the last image bitstream to be encoded are completed and then its quantization is performed.

[0089] By adopting a parallel manner for the processing flow in the first encoding sub-stage, the image encoding efficiency can be improved, achieving an image encoding effect of high density and low latency.

[0090] Step 420: Cache the quantized residual information into the shared macroblock information buffer.

[0091] Among them, the shared macroblock information buffer is connected to the first encoding sub-stage and the second encoding sub-stage, and can realize a complete encoding process when the encoding process is decoupled. The shared macroblock information buffer can support data access of both the first encoding sub-stage and the second encoding sub-stage simultaneously.

[0092] Optionally, caching the quantized residual information into the shared macroblock information buffer includes: performing zigzag scanning on the quantized residual information of a preset number of macroblocks in units of two macroblock rows and caching it into the shared macroblock information buffer.

[0093] Specifically, the Zigzag Buffer can be a buffer for zigzag scanning quantization residual information in units of two macroblock rows. As Figure 5 shown, the Zigzag Buffer completes the data interaction between the first encoding sub-stage (Pipe1) and the second encoding sub-stage (Pipe0). The Zigzag Buffer caches the intermediate macroblock information of the pipe in a zigzag scanning manner, and organizes the data structure of the Zigzag Buffer in units of two macroblock rows (i.e., 32 rows of pixels).

[0094] Through the shared macroblock information buffer, asynchronous and independent processing of the first encoding sub-stage and the second encoding sub-stage can be achieved, thereby decoupling the encoding process, improving the encoding efficiency, and achieving the image encoding effect of high density and low latency.

[0095] Step 430: When performing image encoding, use an asynchronous and independent clock to obtain the quantization residual information to be processed in the macroblock information buffer.

[0096] The first encoding sub-stage and the second encoding sub-stage can be executed asynchronously and independently. Therefore, when performing the second encoding sub-stage processing, an asynchronous and independent clock can be used to obtain the quantization residual information to be processed in the macroblock information buffer.

[0097] Optionally, using an asynchronous and independent clock to obtain the quantization residual information to be processed in the macroblock information buffer includes: using an asynchronous and independent clock to obtain the quantization residual information to be processed in two macroblock rows through zigzag scanning in the macroblock information buffer.

[0098] As Figure 5 shown, the quantization residual information to be processed of MB0, MB1, MB2, MB3, MB4... can be sequentially obtained through zigzag scanning in the macroblock information buffer. Specifically, in the first encoding sub-stage, the quantization residual information is cached in the shared macroblock information buffer, and when obtaining the quantization residual information to be processed in the macroblock information buffer in the second encoding sub-stage, using an asynchronous and independent clock can decouple the image encoding process. The entropy encoding process during image encoding does not need to wait for the quantization process to complete, thereby improving the image encoding efficiency, solving the bottleneck of bit dependence in the code stream during the entropy encoding process, and greatly improving the throughput rate. Moreover, by using an asynchronous and independent clock, the clock frequency of quantization can be effectively increased, and the hardware area efficiency ratio of image encoding can be effectively improved.

[0099] Step 440: Perform the second encoding sub-stage processing on the quantization residual information to be processed to obtain the image encoding code stream.

[0100] Among them, the second encoding sub-stage includes: entropy encoding (EC) and image encoding code stream write-back (SWB).

[0101] Specifically, the quantization residual information to be processed can be obtained from the macroblock information buffer for entropy encoding, and the entropy encoding bitstream is written into the Bit FIFO. The SDMA writes the bitstream in the Bit FIFO into the memory Bitrate buffer to complete the image encoding process.

[0102] To accelerate the image encoding process, when performing the second encoding sub-stage processing on the quantization residual information to be processed, a parallel method can be adopted for the processing flow in the second encoding sub-stage.

[0103] In an optional implementation manner of the embodiment of the present invention, optionally, performing the second encoding sub-stage processing on the quantization residual information to be processed includes: for the quantization residual information to be processed of a preset number of macroblocks, when writing back the image encoding bitstream of the quantization residual information to be processed of the current macroblock, performing entropy encoding on the quantization residual information to be processed of the next macroblock.

[0104] When adopting a parallel method for the processing flow in the second encoding sub-stage, different data volume units can be used. For example, the second encoding sub-stage processing in the parallel method can be performed in units of macroblocks. As Figure 5 shown, entropy encoding and image encoding bitstream write-back can be performed on the quantization residual information to be processed corresponding to MB0; when writing back the image encoding bitstream of the information corresponding to MB0, entropy encoding can be performed on the quantization residual information to be processed corresponding to MB1; then, when writing back the information corresponding to MB1, entropy encoding can be performed on the quantization residual information to be processed corresponding to MB2; and so on, until after the entropy encoding of the last quantization residual information to be processed is completed, its image encoding bitstream write-back is performed.

[0105] By adopting a parallel method for the processing flow in the second encoding sub-stage, the image encoding efficiency can be improved, achieving the image encoding effect of high density and low latency.

[0106] The technical solution of this embodiment is as follows: during image encoding, the bitstream of the image to be encoded is processed in a first encoding sub-stage to obtain quantization residual information. The first encoding sub-stage includes: pixel acquisition of the image to be encoded, discrete cosine transform, and quantization. The quantization residual information is cached in a shared macroblock information buffer. During image encoding, an asynchronous independent clock is used to obtain the quantization residual information to be processed in the macroblock information buffer. The quantization residual information to be processed is processed in a second encoding sub-stage to obtain an image encoding bitstream. The second encoding sub-stage includes: entropy encoding and writing back the image encoding bitstream. This solves the problem of bitstream bit dependence bottleneck during image encoding. By decoupling and splitting the image encoding pipeline and using a shared buffer to interact the bitstream and pixel information, the bitstream bit dependence bottleneck in the entropy encoding process can be effectively solved, and the throughput can be improved. Using an asynchronous independent clock domain in image encoding can effectively increase the clock frequency of discrete cosine transform and quantization, and effectively improve the hardware area efficiency ratio of image encoding.

[0107] Based on the above implementation, in order to reduce the address switching overhead in the image encoding scenario during data interaction and improve the memory utilization rate of the device side, optionally, the method further includes: when performing image encoding data interaction with the host, obtaining a unified virtual address and converting the unified virtual address into a device virtual address; obtaining the register segment mapping bound to the data interaction, and based on the mapping relationship between the virtual address and the physical address in the register segment mapping and the device virtual address, obtaining the device physical address; performing image encoding data interaction through the device physical address.

[0108] Among them, data interaction may refer to reading data from a memory such as a frame buffer or a bit buffer for image encoding. At this time, the address switching overhead can be reduced through the register segment mapping, and the memory utilization rate of the device side can be improved.

[0109] During concurrent multiplexed image encoding (session), each image encoding can independently issue image encoding and decoding tasks at the host side and carry the UVA in the tasks. When the hardware encoding device receives the UVA, it can convert the UVA into a DVA.

[0110] In the embodiment of the present invention, in order to reduce the address switching overhead in the image encoding scenario during data interaction and improve the memory utilization rate of the device side, it can be achieved by, for example Figure 3The register segment mapping shown replaces the traditional MMU page table mapping. According to the mapping relationship between the virtual base address and the physical base address in the register segment mapping, and the address size, the corresponding device physical address can be determined for the device virtual address. Furthermore, the device memory such as the framebuffer or the bit buffer is read and written through the device physical address to complete the image encoding. By replacing the traditional MMU page table mapping with the register segment mapping, the end-to-end latency in the synchronous encoding mode is greatly reduced, the address switching overhead in the synchronous image encoding scenario is reduced, and the device-side memory utilization rate is improved.

[0111] Embodiment III

[0112] Figure 6 FIG. is a flowchart of an image encoding and decoding method provided according to Embodiment III of the present invention. This embodiment is applicable to the situation of improving the encoding and decoding efficiency and throughput during image encoding and decoding. This method can be combined with the above-mentioned image encoding method and / or image decoding method. This method can be executed by an image encoding and decoding device, which can be implemented in the form of hardware and / or software. The image encoding and decoding device can be configured in an electronic device, and the electronic device can be a mobile phone, a tablet computer (PAD), a computer (PC, Personal Computer), etc. As Figure 6 shown, the method includes:

[0113] Step 610, during image decoding, perform a first decoding sub-stage process on the image code stream to be decoded to obtain pixel residual information.

[0114] Among them, the first decoding sub-stage includes: obtaining the code stream of the image to be decoded and entropy decoding.

[0115] Optionally, performing the first decoding sub-stage process on the image code stream to be decoded includes: for the image code stream to be decoded with a preset number of macroblocks, when performing entropy decoding on the image code stream of the current macroblock, obtaining the code stream of the image code stream of the next macroblock.

[0116] Step 620, cache the pixel residual information into a shared macroblock information buffer.

[0117] Optionally, caching the pixel residual information into a shared macroblock information buffer includes: caching the pixel residual information of a preset number of macroblocks into the shared macroblock information buffer in a zigzag scan in units of two macroblock rows.

[0118] Step 630, during image decoding, asynchronously obtain the pixel residual information to be processed in the macroblock information buffer using an independent clock.

[0119] Optionally, an asynchronous independent clock is used to obtain the pixel residual information to be processed in the macroblock information buffer, including: using an asynchronous independent clock to obtain the pixel residual information to be processed in two macroblock rows in the macroblock information buffer through zigzag scanning.

[0120] Step 640: Perform a second decoding sub-stage process on the pixel residual information to be processed to obtain image decoding information.

[0121] Among them, the second decoding sub-stage includes inverse quantization, inverse discrete cosine transform, and pixel bitstream write-back.

[0122] Optionally, performing a second decoding sub-stage process on the pixel residual information to be processed includes: for the pixel residual information to be processed of a preset number of macroblocks, when performing inverse discrete cosine transform on the pixel residual information to be processed of the current macroblock, performing pixel bitstream write-back on the pixel residual information to be processed of the previous macroblock, and performing inverse quantization on the pixel residual information to be processed of the next macroblock.

[0123] Step 650: When encoding an image, perform a first encoding sub-stage process on the image bitstream to be encoded to obtain quantization residual information.

[0124] Among them, the first encoding sub-stage includes: pixel acquisition of the image to be encoded, discrete cosine transform, and quantization.

[0125] Optionally, performing a first encoding sub-stage process on the image bitstream to be encoded includes: for the image bitstream to be encoded of a preset number of macroblocks, when performing discrete cosine transform on the image bitstream to be encoded of the current macroblock, performing quantization on the image bitstream to be encoded of the previous macroblock, and performing pixel acquisition on the image bitstream to be encoded of the next macroblock.

[0126] Step 660: Cache the quantization residual information into the shared macroblock information buffer.

[0127] Optionally, caching the quantization residual information into the shared macroblock information buffer includes: caching the quantization residual information of a preset number of macroblocks into the shared macroblock information buffer by zigzag scanning in units of two macroblock rows.

[0128] Step 670: When encoding an image, use an asynchronous independent clock to obtain the quantization residual information to be processed in the macroblock information buffer.

[0129] Optionally, using an asynchronous independent clock to obtain the quantization residual information to be processed in the macroblock information buffer includes: using an asynchronous independent clock to obtain the quantization residual information to be processed in two macroblock rows in the macroblock information buffer through zigzag scanning.

[0130] Step 680: Perform a second encoding sub-stage process on the quantization residual information to be processed to obtain the image encoding bitstream.

[0131] Among them, the second coding sub-stage includes: entropy coding and writing back of the image coding bitstream.

[0132] Optionally, performing the second coding sub-stage processing on the quantization residual information to be processed includes: for the quantization residual information to be processed of a preset number of macroblocks, when writing back the quantization residual information to be processed of the current macroblock into the image coding bitstream, performing entropy coding on the quantization residual information to be processed of the next macroblock.

[0133] The embodiments of the present invention are not limited to the above execution process. Image coding can be performed first and then image decoding, that is, steps 650 to 680 can be executed first, and then steps 610 to 640 can be executed.

[0134] Figure 7 It is a schematic diagram of an image coding and decoding process provided by Embodiment 3 of the present invention. As Figure 7 shown, in the hardware image decoding data stream, first read the bitstream of the image to be decoded from the bitstream buffer through SDMA into the hardware Bit FIFO, and then read the bitstream from the Bit FIFO for entropy decoding. After the entropy decoding is completed, the pixel residual information is asynchronously written into the Zigzag Buffer and the reconstruction of the pixels is completed. Among them, the Zigzag Buffer can cache the pixel residual information with an image width of 32 rows. Then, an asynchronous independent clock can be used to read the pixel residual information after entropy decoding from the Zigzag Buffer for inverse quantization. After the inverse quantization is completed, inverse discrete cosine transform is performed, and the result is output to the Pixel FIFO. Finally, the VDMA writes the image decoding information back to the memory frame buffer to complete the entire decoding process.

[0135] As Figure 7 shown, in the hardware image coding data stream, first read the pixels of the image to be encoded from the frame buffer through the VDMA module into the hardware Pixel FIFO, and then read the pixels of the image to be encoded from the Pixel FIFO for discrete cosine transform, and then quantization immediately. After the discrete cosine transform and quantization are completed, they can be cached in the Zigzag Buffer in the order of 16×16 macroblocks. Then, entropy coding can be performed on the quantization residual information to be processed in the Zigzag Buffer. Finally, the encoded image coding bitstream is written into the Bit FIFO, and the SDMA writes the image coding bitstream in the bit FIFO into the memory Bitrate buffer to complete the image coding process.

[0136] By decoupling and segmenting the image encoding and decoding pipeline and interacting the bitstream and pixel information through a shared cache, the bitstream bit dependence bottleneck in the entropy encoding and decoding processes can be effectively solved, the throughput can be improved, and the latency in image encoding and decoding can be reduced. It can be widely applied to the AIGC scenario.

[0137] Based on the above embodiments, optionally, the method further includes: when interacting with the host for image decoding and / or encoding data, obtaining a unified virtual address and converting the unified virtual address into a device virtual address; obtaining a register segment mapping bound to the data interaction, and based on the mapping relationship between the virtual address and the physical address in the register segment mapping and the device virtual address, obtaining a device physical address; and performing image decoding and / or encoding data interaction through the device physical address.

[0138] Through dynamic address space mapping, the virtual address space in the image encoding and decoding context is mapped to the device storage physical segment, effectively reducing the address switching overhead in the synchronous image encoding and decoding scenario and improving the memory utilization rate at the device end. In addition, in the address space mapping, a unified address space with different address bit widths can be dynamically adapted without being restricted by the address bit width of the MMU page table entry.

[0139] Embodiment 4

[0140] Figure 8 is a schematic structural diagram of an image decoding device according to Embodiment 4 of the present invention. As Figure 8 shown, the device includes: a first decoding sub-stage processing module 810, a pixel residual information caching module 820, a to-be-processed pixel residual information obtaining module 830, and a second decoding sub-stage processing module 840. Among them:

[0141] The first decoding sub-stage processing module 810 is configured to perform a first decoding sub-stage process on the to-be-decoded image bitstream during image decoding to obtain pixel residual information; wherein, the first decoding sub-stage includes: bitstream acquisition of the to-be-decoded image and entropy decoding;

[0142] The pixel residual information caching module 820 is configured to cache the pixel residual information into a shared macroblock information cache;

[0143] The to-be-processed pixel residual information obtaining module 830 is configured to obtain the to-be-processed pixel residual information in the macroblock information cache using an asynchronous independent clock during image decoding;

[0144] The second decoding sub-stage processing module 840 is configured to perform a second decoding sub-stage process on the to-be-processed pixel residual information to obtain image decoding information; wherein, the second decoding sub-stage includes inverse quantization, inverse discrete cosine transform, and pixel bitstream write-back.

[0145] Optionally, the first decoding sub-stage processing module 810 is specifically configured to: when entropy decoding the image code stream of a preset number of macroblocks, obtain the code stream of the image code stream of the next macroblock when entropy decoding the image code stream of the current macroblock.

[0146] Optionally, the pixel residual information caching module 820 is specifically configured to: cache the pixel residual information of a preset number of macroblocks into the shared macroblock information cache in a zigzag scan in units of two macroblock rows.

[0147] Optionally, the pixel residual information to be processed obtaining module 830 is specifically configured to: obtain the pixel residual information to be processed in two macroblock rows in the macroblock information cache through a zigzag scan using an asynchronous independent clock.

[0148] Optionally, the second decoding sub-stage processing module 840 is specifically configured to: for the pixel residual information to be processed of a preset number of macroblocks, when performing inverse discrete cosine transform on the pixel residual information to be processed of the current macroblock, write back the pixel code stream of the pixel residual information to be processed of the previous macroblock, and perform inverse quantization on the pixel residual information to be processed of the next macroblock.

[0149] Based on the above embodiments, optionally, the image decoding device further includes:

[0150] A device virtual address determination module, configured to obtain a unified virtual address when interacting with the host for image decoding and / or encoding data, and convert the unified virtual address into a device virtual address;

[0151] A device physical address determination module, configured to obtain a register segment mapping bound to the data interaction, and obtain the device physical address according to the mapping relationship between the virtual address and the physical address in the register segment mapping, and the device virtual address;

[0152] A data interaction module, configured to perform image decoding and / or encoding data interaction through the device physical address.

[0153] The image decoding device provided by the embodiments of the present invention can execute the image decoding method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0154] Embodiment Five

[0155] Figure 9 It is a schematic structural diagram of an image encoding device according to Embodiment Five of the present invention. As Figure 9 shown, the device includes: a first encoding sub-stage processing module 910, a quantization residual information caching module 920, a quantization residual information to be processed obtaining module 930, and a second encoding sub-stage processing module 940. Among them:

[0156] The first encoding sub-stage processing module 910 is used to perform the first encoding sub-stage processing on the image bitstream to be encoded during image encoding, so as to obtain quantization residual information; wherein, the first encoding sub-stage includes: pixel acquisition of the image to be encoded, discrete cosine transform, and quantization.

[0157] The quantization residual information caching module 920 is used to cache the quantization residual information into the shared macroblock information cache.

[0158] The quantization residual information to be processed acquisition module 930 is used to acquire the quantization residual information to be processed in the macroblock information cache by using an asynchronous independent clock during image encoding.

[0159] The second encoding sub-stage processing module 940 is used to perform the second encoding sub-stage processing on the quantization residual information to be processed, so as to obtain the image encoding bitstream; wherein, the second encoding sub-stage includes: entropy encoding and writing back of the image encoding bitstream.

[0160] Optionally, the first encoding sub-stage processing module 910 is specifically used for: for the image bitstream to be encoded of a preset number of macroblocks, when performing discrete cosine transform on the image bitstream to be encoded of the current macroblock, performing quantization on the image bitstream to be encoded of the previous macroblock, and performing pixel acquisition on the image bitstream to be encoded of the next macroblock.

[0161] Optionally, the quantization residual information caching module 920 is specifically used for: caching the quantization residual information of a preset number of macroblocks into the shared macroblock information cache by performing zigzag scan in units of two macroblock rows.

[0162] Optionally, the quantization residual information to be processed acquisition module 930 is specifically used for: acquiring the quantization residual information to be processed in two macroblock rows in the macroblock information cache by performing zigzag scan by using an asynchronous independent clock.

[0163] Optionally, the second encoding sub-stage processing module 940 is specifically used for: for the quantization residual information to be processed of a preset number of macroblocks, when writing back the image encoding bitstream of the quantization residual information to be processed of the current macroblock, performing entropy encoding on the quantization residual information to be processed of the next macroblock.

[0164] Based on the above embodiments, optionally, the image encoding device further includes:

[0165] The device virtual address determination module is used to obtain a unified virtual address and convert the unified virtual address into a device virtual address when performing image encoding data interaction with the host.

[0166] The device physical address determination module is used to obtain the register segment mapping bound to data interaction, and based on the mapping relationship between the virtual address and the physical address in the register segment mapping, and the device virtual address, obtain the device physical address;

[0167] The data interaction module is used to perform image encoding data interaction through the device physical address.

[0168] The image encoding device provided by the embodiments of the present invention can execute the image encoding method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0169] Embodiment Six

[0170] Figure 10 is a structural schematic diagram of an image encoding and decoding device provided by Embodiment Six of the present invention. As Figure 10 shown, the device includes: a first decoding sub-stage processing module 1010, a pixel residual information caching module 1020, a pixel residual information to be processed acquisition module 1030, a second decoding sub-stage processing module 1040, a first encoding sub-stage processing module 1050, a quantization residual information caching module 1060, a quantization residual information to be processed acquisition module 1070, and a second encoding sub-stage processing module 1080. Among them:

[0171] The first decoding sub-stage processing module 1010 is used to perform the first decoding sub-stage processing on the image code stream to be decoded during image decoding to obtain pixel residual information; wherein, the first decoding sub-stage includes: code stream acquisition of the image to be decoded and entropy decoding;

[0172] The pixel residual information caching module 1020 is used to cache the pixel residual information into a shared macroblock information cache;

[0173] The pixel residual information to be processed acquisition module 1030 is used to asynchronously and independently acquire the pixel residual information to be processed in the macroblock information cache during image decoding;

[0174] The second decoding sub-stage processing module 1040 is used to perform the second decoding sub-stage processing on the pixel residual information to be processed to obtain image decoding information; wherein, the second decoding sub-stage includes: inverse quantization, inverse discrete cosine transform, and pixel code stream write-back;

[0175] The first encoding sub-stage processing module 1050 is used to perform the first encoding sub-stage processing on the image code stream to be encoded during image encoding to obtain quantization residual information; wherein, the first encoding sub-stage includes: pixel acquisition of the image to be encoded, discrete cosine transform, and quantization;

[0176] The quantization residual information caching module 1060 is configured to cache quantization residual information into a shared macroblock information cache;

[0177] The quantization residual information to be processed acquisition module 1070 is configured to acquire quantization residual information to be processed in the macroblock information cache by using an asynchronous independent clock during image encoding;

[0178] The second encoding sub-stage processing module 1080 is configured to perform second encoding sub-stage processing on the quantization residual information to be processed to obtain an image encoding bitstream; wherein, the second encoding sub-stage includes: entropy encoding and writing back the image encoding bitstream.

[0179] The image encoding and decoding apparatus provided by the embodiments of the present invention can execute the image encoding and decoding method provided by any embodiment of the present invention, and has functional modules and beneficial effects corresponding to the execution of the method.

[0180] Embodiment Seven

[0181] Figure 11 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described herein and / or claimed.

[0182] As Figure 11 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein, the memory stores a computer program executable by at least one processor, and the processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0183] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0184] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as an image decoding method and / or an image encoding method.

[0185] In some embodiments, the image decoding method and / or the image encoding method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the image decoding method and / or the image encoding method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the image decoding method and / or the image encoding method by any other suitable means (e.g., by means of firmware).

[0186] The various embodiments of the systems and technologies described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0187] A computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing device, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0188] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0189] In order to provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0190] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected with each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0191] A computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs that run on respective computers and have a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0192] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.

[0193] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. An image decoding method, characterized in that: include: During image decoding, the first decoding sub-stage is performed on the code stream of the image to be decoded to obtain pixel residual information; wherein the first decoding sub-stage includes: obtaining the code stream of the image to be decoded and entropy decoding; Buffering the pixel residual information in a shared macroblock information buffer; When decoding an image, an asynchronous independent clock is used to obtain pixel residual information to be processed in the macroblock information buffer; The pixel residual information to be processed is processed in a second decoding sub-stage to obtain image decoding information; wherein the second decoding sub-stage includes inverse quantization, inverse discrete cosine transform and pixel code stream write-back.

2. The method according to claim 1, characterized in that The first decoding sub-stage processing is performed on the image code stream to be decoded, including: For the to-be-decoded image code streams of a preset number of macroblocks, when entropy decoding is performed on the to-be-decoded image code stream of the current macroblock, code stream acquisition is performed on the to-be-decoded image code stream of the next macroblock.

3. The method according to claim 1, characterized in that The second decoding sub-stage processing is performed on the pixel residual information to be processed, including: For the preset number of macroblocks of pixel residual information to be processed, when the pixel residual information to be processed of the current macroblock is subjected to inverse discrete cosine transform, the pixel residual information to be processed of the previous macroblock is written back to the pixel code stream, and the pixel residual information to be processed of the next macroblock is inversely quantized.

4. The method according to claim 1, characterized in that: Cache the pixel residual information in a shared macroblock information buffer, including: Buffering the pixel residual information of a preset number of macroblocks in a zigzag pattern in units of two macroblock rows into a shared macroblock information buffer; Acquiring the pixel residual information to be processed in the macroblock information buffer using an asynchronous independent clock, including: An asynchronous independent clock is used to obtain the to-be-processed pixel residual information in two macroblock rows through zigzag scanning in the macroblock information buffer.

5. The method according to claim 1, characterized in that Also includes: During image encoding, a first encoding sub-stage is performed on a code stream of an image to be encoded to obtain quantized residual information; wherein the first encoding sub-stage includes: pixel acquisition, discrete cosine transform and quantization of the image to be encoded; Buffering the quantized residual information in a shared macroblock information buffer; During image encoding, an asynchronous independent clock is used to obtain the quantized residual information to be processed in the macroblock information buffer; The quantized residual information to be processed is processed in a second encoding sub-stage to obtain an image encoding code stream; wherein the second encoding sub-stage includes: entropy encoding and writing back the image encoding code stream.

6. The method according to any one of claims 1 to 5, characterized in that: Also includes: When performing image decoding and / or encoding data interaction with the host, obtaining a unified virtual address and converting the unified virtual address into a device virtual address; Acquire a register segment mapping bound to the data interaction, and obtain a device physical address according to a mapping relationship between a virtual address and a physical address in the register segment mapping and the device virtual address; Image decoding and / or encoding data interaction is performed through the device physical address.

7. An image encoding method, characterized in that: include: During image encoding, a first encoding sub-stage is performed on a code stream of an image to be encoded to obtain quantized residual information; wherein the first encoding sub-stage includes: pixel acquisition, discrete cosine transform and quantization of the image to be encoded; Buffering the quantized residual information in a shared macroblock information buffer; During image encoding, an asynchronous independent clock is used to obtain the quantized residual information to be processed in the macroblock information buffer; The quantized residual information to be processed is processed in a second encoding sub-stage to obtain an image encoding code stream; wherein the second encoding sub-stage includes: entropy encoding and writing back the image encoding code stream.

8. The method according to claim 7, characterized in that Also includes: When interacting with the host for image coding data, obtaining a unified virtual address and converting the unified virtual address into a device virtual address; Acquire a register segment mapping bound to the data interaction, and obtain a device physical address according to a mapping relationship between a virtual address and a physical address in the register segment mapping and the device virtual address; Image encoding data is exchanged via the device physical address.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the image decoding method described in any one of claims 1-6, or can execute the image encoding method described in any one of claims 7-8.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the image decoding method described in any one of claims 1 to 6, or to implement the image encoding method described in any one of claims 7 to 8 when executed.

Citation Information

Cited By

  • Image decoding heterogeneous acceleration system architecture based on CPU + FPGA

    CN121691713A