Image compression method and image compression architecture
Patent Information
- Application Number
- CN202610353652.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-23
- Publication Date
- 2026-08-18
- Estimated Expiration
- 2046-03-23
AI Technical Summary
[0003]相关技术中,片外访存单元和片内计算单元之间的图像压缩算法所需计算和存储需求大,使得片外访存单元在对图像进行压缩的过程中,需要消耗过多资源,导致压缩效率低
[0015] According to embodiments of this application, by dividing the image to be compressed into multiple sub-band images, the image compression architecture can capture image features at different resolution levels during the compression process. Furthermore, by dividing the multiple sub-band images into blocks, the amount of data processed by the image compression architecture in a single operation is reduced, thus avoiding excessive computational and storage resources. In addition, by utilizing the input channel layer corresponding to the image block in the image compression model to perform feature fusion on the multiple sub-band images and the reference image block, complementary feature information provided by different sub-band images can be combined. Moreover, the target image is constructed on-chip based on the image fusion features corresponding to each of the multiple image blocks, without interaction with off-chip access units. This avoids the problem of low processing efficiency caused by long interaction times, thereby improving processing efficiency while avoiding excessive consumption of computational and storage resources.
Smart Images

Figure CN121883622B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, specifically to the field of image compression technology, and more specifically to an image compression method and an image compression architecture. Background Technology
[0002] Image compression algorithms are used to compress the amount of data in an image to obtain a compressed image with a smaller data size.
[0003] In related technologies, the image compression algorithm between off-chip memory access units and on-chip computing units requires a large amount of computation and storage, which causes the off-chip memory access unit to consume too many resources during the image compression process, resulting in low compression efficiency. Summary of the Invention
[0004] In view of the above problems, this application provides an image compression method and an image compression architecture.
[0005] According to a first aspect of this application, an image compression method is provided, applied to an image compression architecture, the image compression architecture including an off-chip memory access unit and an on-chip computing unit, comprising: receiving an image to be compressed provided by the off-chip memory access unit using the on-chip computing unit, and performing the following operations using the on-chip computing unit: dividing multiple sub-band images of the image to be compressed into blocks, obtaining multiple image blocks for each of the multiple sub-band images, the multiple sub-band images corresponding to the same block region, and the multiple sub-band images having different image resolutions; performing feature fusion on the multiple sub-band images corresponding to the image blocks and a reference image block using the input channel layer corresponding to the image blocks in the image compression model, obtaining image fusion features corresponding to the image blocks, wherein the reference image block is determined by fusing multiple target image blocks; constructing a target image by fusing the image fusion features corresponding to each of the multiple image blocks, the data size of the target image being smaller than the data size of the image to be compressed; and sending the target image to the off-chip memory access unit.
[0006] According to embodiments of this application, feature fusion is performed on multiple sub-band images and reference image blocks corresponding to an image block using the input channel layer corresponding to the image block in the image compression model to obtain image fusion features corresponding to the image block. This includes: for any sub-band image among the multiple sub-band images, feature extraction is performed on the sub-band image and reference image block corresponding to the image block using the input channel layer corresponding to the image block in the image compression model, respectively, to obtain sub-band image features of the sub-band image and reference image features of the reference image block; the sub-band image features and reference image features are fused to obtain intermediate fused image features; feature extraction is performed on the intermediate fused image features to obtain image fusion features of the sub-band image; and image fusion features corresponding to the image block are obtained based on the image fusion features of each of the multiple sub-band images.
[0007] According to an embodiment of this application, feature extraction is performed on the sub-band image and reference image block corresponding to the image block using the input channel layer corresponding to the image block in the image compression model, to obtain the sub-band image features of the sub-band image and the reference image features of the reference image block. This includes: extracting features from the sub-band image and the reference image block based on the first input channel weight of the target input channel and the first output channel weight of the target output channel corresponding to the target input channel in the input channel layer, to obtain the sub-band image features of the sub-band image and the reference image features of the reference image block.
[0008] According to an embodiment of this application, feature extraction is performed on the sub-band image to obtain the sub-band image features of the sub-band image, including: for the nth convolutional kernel position among the N convolutional kernel positions corresponding to the input channel layer, feature extraction is performed on the sub-band image and the (n-1)th convolutional kernel position features based on the first input channel weight, the first output channel weight, and the nth convolutional kernel position to obtain the nth convolutional kernel position features; when n=N, the nth convolutional kernel position features are determined as the sub-band image features of the sub-band image; when n=1, feature extraction is performed on the sub-band image based on the first input channel weight, the first output channel weight, and the first convolutional kernel position to obtain the first convolutional kernel position features, 1≤n≤N, where n and N are both integers.
[0009] According to an embodiment of this application, feature extraction is performed on the intermediate fused image features to obtain the image fusion features of the sub-band image, including: using the second input channel weight of the target fusion input channel of the fusion input channel layer corresponding to the intermediate fused image features in the image compression model and the second output channel weight of the target fusion output channel corresponding to the target fusion input channel to extract features from the intermediate fused image features to obtain the image fusion features of the sub-band image.
[0010] According to an embodiment of this application, feature extraction is performed on the intermediate fused image features using the second input channel weight of the target fused input channel layer corresponding to the intermediate fused image features in the image compression model and the second output channel weight of the target fused output channel corresponding to the target fused input channel, to obtain the image fusion features of the sub-band image. This includes: when there are multiple target fused input channels, performing the following operations: for the m-th target fused input channel among the M target fused input channels, based on the m-th second input channel weight and the m-th second output channel weight, extracting the m-th intermediate fused image channel features and the (m-1)-th image fusion features indicated by the m-th target fused input channel. Feature extraction is performed to obtain the image fusion feature of the m-th image. When m=M, the image fusion feature of the m-th image is determined as the image fusion feature of the sub-band image indicated by multiple target fusion input channels. When m=1, based on the weights of the first second input channel and the weights of the first second output channel, feature extraction is performed on the feature of the first intermediate fusion image channel indicated by the first target fusion input channel to obtain the first image fusion feature. Here, the weights of the M second input channels are different, the weights of the M second output channels are different, 1≤m≤M, and m and M are both integers. Based on the image fusion feature of the sub-band image corresponding to the target fusion input channel of the fusion input channel layer, the image fusion feature of the sub-band image is obtained.
[0011] According to an embodiment of this application, the compression architecture further includes an on-chip storage unit; the method further includes: after obtaining the target image features corresponding to the target fusion input channel, clearing the intermediate data corresponding to the target fusion input channel, and storing the target image features corresponding to the target fusion input channel in the on-chip storage unit, wherein the intermediate data is the data generated during the process of generating the target image features corresponding to the target fusion input channel.
[0012] According to an embodiment of this application, the image to be compressed is divided into multiple sub-band images to obtain multiple image blocks for each sub-band image. This includes: analyzing the bandwidth and storage requirements of a preset initial block area based on the resource data of the image compression architecture, obtaining analysis results, whereby the analysis results characterize whether the resource data of the image compression architecture can provide the bandwidth resources indicated by the bandwidth requirements of the initial block area and the storage resources indicated by the storage requirements of the initial block area. The resource data is the resource data used by the image compression architecture to process the image to be compressed; optimizing the size of the initial block area based on the analysis results to obtain a target block area; and dividing the image to be compressed into multiple sub-band images based on the target block area to obtain multiple image blocks for each sub-band image.
[0013] According to an embodiment of this application, the above method further includes: performing discrete wavelet transform on the image to be compressed to obtain images to be compressed with different image resolutions; and using the images to be compressed with different image resolutions as multiple sub-band images of the image to be compressed.
[0014] A second aspect of this application provides an image compression architecture, comprising: an off-chip memory access unit for storing an image to be compressed and receiving a compressed target image; and an on-chip computing unit for performing the following operations: dividing multiple sub-band images of the image to be compressed into blocks to obtain multiple image blocks for each of the multiple sub-band images, wherein the multiple sub-band images correspond to the same block region and the multiple sub-band images have different image resolutions; using the input channel layer corresponding to the image block in the image compression model to perform feature fusion on the multiple sub-band images corresponding to the image block and a reference image block to obtain image fusion features corresponding to the image block, wherein the reference image block is determined by fusing multiple target image blocks; constructing a target image by fusing the image fusion features corresponding to each of the multiple image blocks, wherein the data size of the target image is smaller than the data size of the image to be compressed; and sending the target image to the off-chip memory access unit.
[0015] According to embodiments of this application, by dividing the image to be compressed into multiple sub-band images, the image compression architecture can capture image features at different resolution levels during the compression process. Furthermore, by dividing the multiple sub-band images into blocks, the amount of data processed by the image compression architecture in a single operation is reduced, thus avoiding excessive computational and storage resources. In addition, by utilizing the input channel layer corresponding to the image block in the image compression model to perform feature fusion on the multiple sub-band images and the reference image block, complementary feature information provided by different sub-band images can be combined. Moreover, the target image is constructed on-chip based on the image fusion features corresponding to each of the multiple image blocks, without interaction with off-chip access units. This avoids the problem of low processing efficiency caused by long interaction times, thereby improving processing efficiency while avoiding excessive consumption of computational and storage resources. Attached Figure Description
[0016] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0017] Figure 1 An application scenario diagram of the image compression method according to an embodiment of this application is shown;
[0018] Figure 2 A flowchart of an image compression method according to an embodiment of this application is shown;
[0019] Figure 3A schematic diagram of an image compression model according to an embodiment of this application is shown;
[0020] Figure 4 This illustration shows a schematic diagram of feature extraction performed on a sub-band image and a reference image block corresponding to an image block, according to an embodiment of this application.
[0021] Figure 5 This illustration shows a schematic diagram of image fusion features of a sub-band image obtained by feature extraction of intermediate fused image features according to an embodiment of this application;
[0022] Figure 6 A schematic diagram illustrating feature extraction of a sub-band image according to an embodiment of this application is shown;
[0023] Figure 7 A block diagram illustrating the relevant technology according to embodiments of this application is shown;
[0024] Figure 8 A block diagram according to an embodiment of this application is shown;
[0025] Figure 9 A schematic diagram of performing discrete wavelet transform on an image to be compressed according to an embodiment of this application is shown;
[0026] Figure 10 A schematic diagram of a sub-band image according to an embodiment of this application is shown;
[0027] Figure 11 An image compression architecture according to an embodiment of this application is shown. Detailed Implementation
[0028] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.
[0029] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0030] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0031] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0032] Image compression algorithms have unique network architectures: high-resolution input images (4K and above), image size does not decrease with increasing network depth, and the number of parameters is large and the network structure is complex (containing many residual connections and skip structures). These characteristics lead to significant storage and computational requirements.
[0033] End-to-end image compression algorithms include floating-point inference based on Graphics Processing Units (GPUs) and fixed-point inference based on dedicated convolutional neural network accelerators. However, these methods are not suitable for edge applications with strict power consumption requirements due to their high power consumption; furthermore, accelerators cannot fully meet the computational and storage demands of image compression algorithms for high-resolution image inference and complex network structures.
[0034] Accelerators operate on a layer-by-layer inference model, meaning that the output of each layer needs to be stored in off-chip memory (Dynamic Random Access Memory, DRAM) and then loaded again from off-chip DRAM for the next layer. This model results in a large number of DRAM access operations, which becomes a major factor affecting computation and power consumption, especially when dealing with high-definition image output (4K and above). Some accelerators use layer fusion techniques to avoid DRAM access in intermediate layers, but this still requires a significant amount of on-chip storage space to store overlapping data. Furthermore, the phenomenon that feature map size decreases with layer depth during layer fusion contradicts the feature map size invariance in end-to-end image compression networks, and the high-resolution requirements limit its application in end-to-end image compression algorithms.
[0035] In view of this, embodiments of this application provide an image compression method applied to an image compression architecture, the image compression architecture including an off-chip memory access unit and an on-chip computing unit, comprising: receiving an image to be compressed provided by the off-chip memory access unit using the on-chip computing unit, and performing the following operations using the on-chip computing unit: dividing multiple sub-band images of the image to be compressed into blocks, obtaining multiple image blocks for each of the multiple sub-band images, the multiple sub-band images corresponding to the same block region, the multiple sub-band images having different image resolutions; performing feature fusion on the multiple sub-band images and a reference image block corresponding to the image block using the input channel layer corresponding to the image block in the image compression model, obtaining image fusion features corresponding to the image block, wherein the reference image block is determined by fusing multiple target image blocks; constructing a target image by fusing the image fusion features corresponding to each of the multiple image blocks, the data size of the target image being smaller than the data size of the image to be compressed; and sending the target image to the off-chip memory access unit.
[0036] Figure 1 An application scenario diagram of the image compression method according to an embodiment of this application is shown.
[0037] like Figure 1 As shown, the application scenarios according to this embodiment may include off-chip memory access units and on-chip computing units.
[0038] The off-chip memory access unit is used to provide the image to be compressed to the on-chip computing unit and to receive the corresponding target image.
[0039] The on-chip computing unit is used to receive the image to be compressed, encode the image to be compressed, and further input it into the image compression model for processing, so as to decode the processing result and construct the target image based on the decoded processing result.
[0040] The following will be based on Figure 1 The described scene, through Figures 2-10 The image compression method of the application embodiments will be described in detail.
[0041] Figure 2 A flowchart of an image compression method according to an embodiment of this application is shown.
[0042] Image compression architecture includes off-chip memory access units and on-chip computing units.
[0043] The on-chip computing unit receives the image to be compressed provided by the off-chip memory access unit, and performs operations such as... Figure 2 Operations S210 to S240 are shown.
[0044] In operation S210, the multiple sub-band images of the image to be compressed are divided into blocks, resulting in multiple image blocks for each of the multiple sub-band images.
[0045] Multiple sub-band images correspond to the same block region, and the multiple sub-band images have different image resolutions.
[0046] In operation S220, the input channel layer corresponding to the image block in the image compression model is used to perform feature fusion on multiple sub-band images and reference image blocks corresponding to the image block to obtain the image fusion features corresponding to the image block.
[0047] The reference image block is determined by fusing multiple target image blocks.
[0048] In operation S230, the target image is constructed by fusing the image fusion features corresponding to multiple image blocks.
[0049] The data size of the target image is smaller than that of the image to be compressed.
[0050] In operation S240, the target image is sent to the off-chip access unit.
[0051] The off-chip memory access unit can be DRAM (Dynamic Random Access Memory), used to provide the image to be compressed to the on-chip computing unit and receive the corresponding target image.
[0052] The on-chip computing unit may include multiple image input subunits, multiple computing subunits, multiple fusion subunits, and multiple interaction subunits, wherein the multiple image input subunits are used to perform the above operation S210; the multiple computing subunits are used to perform the above operation S220; the multiple fusion subunits are used to perform the above operation S230; and the multiple interaction subunits are used to perform the above operation S240.
[0053] Multiple sub-band images of the image to be compressed can be multiple images with different image resolutions obtained by dividing the image to be compressed into sub-bands.
[0054] Multiple image blocks for each of the multiple sub-band images can be achieved by physically segmenting the sub-band images into multiple image blocks. The size of the multiple image blocks can be determined based on the resource data of the image compression architecture, so that the resource data of the image compression architecture can support the processing of image blocks of different sizes.
[0055] The image compression model can be an entropy model including multiple convolutional layers, used to extract features from input data and output feature extraction results. For the image segmentation of the embodiments of this application, the image compression model can extract features from multiple sub-band images and reference image blocks corresponding to the image segment to output the result used to construct the target image.
[0056] In an image compression model, the input channel layer corresponding to an image block can be multiple first input channels used to input the image block into the image compression model. These multiple first input channels have different first input channel weights, so that the image block can be processed based on the different first input channel weights.
[0057] The reference image block can be determined by fusing multiple target image blocks. Specifically, multiple sub-band images corresponding to an image block are sorted in ascending order according to their respective image resolutions. For any sub-band image among these sub-band images, multiple preceding sub-band images are determined based on the sorting result, and these preceding sub-band images are fused to obtain the reference image block for that sub-band image. For example, if there are four sub-band images, and these four sub-band images are sorted according to their respective image resolutions as: sub-band image A, sub-band image B, sub-band image C, and sub-band image D, then the reference image block for sub-band image C is obtained by fusing sub-band images A and B.
[0058] The input channel layer corresponding to the image block in the image compression model is used to perform feature fusion on multiple sub-band images and reference image blocks corresponding to the image block. This can be done by extracting features and fusing features on the sub-band images and reference images respectively based on the first input channel weight indicated by the input channel layer, so as to obtain the image fusion features of each of the multiple sub-band images, and then determining the image fusion features corresponding to the image block based on the image fusion features of each of the multiple sub-band images.
[0059] The image fusion features corresponding to multiple image blocks are fused to determine the image fusion features of the image to be compressed. Furthermore, a target image is constructed based on the image fusion features of the image to be compressed, which can be a compressed image of the image to be compressed.
[0060] After the on-chip computing unit obtains the target image, it sends the target image to the off-chip memory access unit.
[0061] According to embodiments of this application, by dividing the image to be compressed into multiple sub-band images, the image compression architecture can capture image features at different resolution levels during the compression process. Furthermore, by dividing the multiple sub-band images into blocks, the amount of data processed by the image compression architecture in a single operation is reduced, thus avoiding excessive computational and storage resources. In addition, by utilizing the input channel layer corresponding to the image block in the image compression model to perform feature fusion on the multiple sub-band images and the reference image block, complementary feature information provided by different sub-band images can be combined. Moreover, the target image is constructed on-chip based on the image fusion features corresponding to each of the multiple image blocks, without interaction with off-chip access units. This avoids the problem of low processing efficiency caused by long interaction times, thereby improving processing efficiency while avoiding excessive consumption of computational and storage resources.
[0062] According to embodiments of this application, feature fusion is performed on multiple sub-band images and reference image blocks corresponding to an image block using the input channel layer corresponding to the image block in the image compression model to obtain image fusion features corresponding to the image block. This includes: for any sub-band image among the multiple sub-band images, feature extraction is performed on the sub-band image and reference image block corresponding to the image block using the input channel layer corresponding to the image block in the image compression model, respectively, to obtain sub-band image features of the sub-band image and reference image features of the reference image block; the sub-band image features and reference image features are fused to obtain intermediate fused image features; feature extraction is performed on the intermediate fused image features to obtain image fusion features of the sub-band image; and image fusion features corresponding to the image block are obtained based on the image fusion features of each of the multiple sub-band images.
[0063] Figure 3 A schematic diagram of an image compression model according to an embodiment of this application is shown.
[0064] like Figure 3 As shown, based on Figure 3 The current information extraction network shown utilizes the first input channel weights indicated by the input channel layer corresponding to the image block in the image compression model to extract features from the sub-band image, obtaining the sub-band image features; based on Figure 3 The reference information extraction network shown uses the first input channel weights indicated by the input channel layer corresponding to the image block in the image compression model to extract features from the reference image and obtain the reference image features of the reference image block.
[0065] The current information extraction network includes an input layer “subband” and multiple convolutional layers “Conv_pre” and “Conv_pre2”; the reference information extraction network includes an input layer “context”, an upsampling layer “upsampling”, and multiple convolutional layers “Conv_c” and “Conv_c2”.
[0066] After obtaining the sub-band image features and reference image features, using Figure 3 The fusion layer “Add” shown fuses the features of the sub-band image and the reference image to obtain intermediate fused image features.
[0067] After obtaining the intermediate fused image features, use Figure 3 The multiple convolutional layers “Conv1”, “Conv2” and “Conv_post” shown extract features from the intermediate fused image to obtain the image fusion features of the sub-band image.
[0068] After determining the image fusion features of each of the multiple sub-band images, the image fusion features of each of the multiple sub-band images are fused to obtain the image fusion features corresponding to the image blocks.
[0069] According to embodiments of this application, by fusing sub-band image features and reference image features, image features are captured at different resolution levels. The complementary feature information provided by different resolution levels improves the accuracy of image fusion features when extracting features from intermediate fused image features.
[0070] According to an embodiment of this application, feature extraction is performed on the sub-band image and reference image block corresponding to the image block using the input channel layer corresponding to the image block in the image compression model, to obtain the sub-band image features of the sub-band image and the reference image features of the reference image block. This includes: extracting features from the sub-band image and the reference image block based on the first input channel weight of the target input channel and the first output channel weight of the target output channel corresponding to the target input channel in the input channel layer, to obtain the sub-band image features of the sub-band image and the reference image features of the reference image block.
[0071] The input channel layer includes multiple first input channels, each with a different weight. These multiple first input channels correspond to multiple first output channels, and the number of each type of first input channel and first output channel is selected based on actual needs. For example, Figure 3 In the current information extraction network shown, "Conv_pre" has 128 first input channels and 1 first output channel, and "Conv_pre2" has 128 first input channels and 128 first output channels.
[0072] In the process of processing image blocks using an image compression model, multiple image blocks indicated by multiple first input channels can be processed sequentially, and multiple first input channels can be grouped, with each group including at least one first input channel.
[0073] Based on the first input channel weight of the target input channel and the first output channel weight of the target output channel corresponding to the target input channel in the input channel layer, feature extraction is performed on the sub-band image. This can be done in the case of only one first input channel and one first output channel, by extracting features from the sub-band image based on the first input channel weight and the first output channel weight, and outputting the sub-band image features corresponding to the first output channel.
[0074] Alternatively, in the case of multiple first input channels and multiple first output channels, for the k-th first input channel and the h-th first output channel out of the K first input channels and H first output channels, feature extraction is performed on the sub-band image and the (k-1)-th sub-band image features using the weights of the k-th first input channel and the h-th first output channel, to obtain the k-th sub-band image features corresponding to the h-th first output channel. This operation is repeated until K sub-band image features corresponding to each of the H first output channels are output.
[0075] After determining the multiple sub-band image features corresponding to each of the multiple first output channels, for any one of the multiple first output channels, the multiple sub-band image features corresponding to that first output channel are fused to determine the fused sub-band image features corresponding to each of the multiple first output channels as the sub-band image features of the sub-band image.
[0076] Similar to the operation of determining the sub-band image features, feature extraction is performed on the reference image features based on the above operation to obtain the reference image features of the reference image blocks.
[0077] According to the embodiments of this application, since a specified number of first output channels and image blocks corresponding to the first output channels are processed sequentially, the amount of data processed in a single process is smaller than the amount of data processed without splitting the channels, thereby improving the efficiency of processing image blocks.
[0078] According to an embodiment of this application, feature extraction is performed on the sub-band image to obtain the sub-band image features of the sub-band image, including: for the nth convolutional kernel position among the N convolutional kernel positions corresponding to the input channel layer, feature extraction is performed on the sub-band image and the (n-1)th convolutional kernel position features based on the first input channel weight, the first output channel weight, and the nth convolutional kernel position to obtain the nth convolutional kernel position features; when n=N, the nth convolutional kernel position features are determined as the sub-band image features of the sub-band image; when n=1, feature extraction is performed on the sub-band image based on the first input channel weight, the first output channel weight, and the first convolutional kernel position to obtain the first convolutional kernel position features, 1≤n≤N, where n and N are both integers.
[0079] Different convolution kernel positions are used to extract features from different locations in the sub-band image.
[0080] For example, Figure 3 In the current information extraction network shown, "Conv_pre" is The convolution kernel has 9 kernel positions, and "Conv_pre2" is... The convolution kernel has 9 kernel positions.
[0081] Figure 4 This illustration shows a schematic diagram of feature extraction performed on a sub-band image and a reference image block corresponding to an image block, according to an embodiment of this application.
[0082] like Figure 4 As shown, there are two first input channels and two first output channels, and Taking the convolution kernel as an example, Ibuf0-Ibuf7 are all sub-band images; W and G indicate two first output channels. For W0 and G0, they respectively indicate the weights of the 0th first input channel and the Wth first output channel, and the weights of the 0th first input channel and the Gth first output channel; for W1 and G1, they respectively indicate the weights of the 1st first input channel and the Wth first output channel, and the weights of the 1st first input channel and the Gth first output channel; (0,0), (0,1), (1,0) and (1,1) correspond to The convolution kernel indicates the four convolution kernel positions.
[0083] like Figure 4As shown, W0(0,0) is used to extract features from the sub-band image output by Ibuf0 based on the weights of the 0th first input channel, the Wth first output channel, and the (0,0)th convolutional kernel position to obtain the first convolutional kernel position feature; W0(0,1) is used to extract features from the sub-band image output by Ibuf1 and the first convolutional kernel position feature based on the weights of the 0th first input channel, the Wth first output channel, and the (0,1)th convolutional kernel position to obtain the second convolutional kernel position feature.
[0084] G0(0,0) is used to extract features from the sub-band image output by Ibuf0 based on the weights of the 0th first input channel, the Gth first output channel, and the (0,0)th convolutional kernel position to obtain the first convolutional kernel position feature; G0(0,1) is used to extract features from the sub-band image output by Ibuf1 and the first convolutional kernel position feature based on the weights of the 0th first input channel, the Gth first output channel, and the (0,1)th convolutional kernel position to obtain the second convolutional kernel position feature.
[0085] According to embodiments of this application, by performing feature extraction based on multiple convolutional kernel locations, it is possible to capture local details and global information of convolutional kernels of different sizes, thereby improving the accuracy of feature extraction.
[0086] According to an embodiment of this application, feature extraction is performed on the intermediate fused image features to obtain the image fusion features of the sub-band image, including: using the second input channel weight of the target fusion input channel of the fusion input channel layer corresponding to the intermediate fused image features in the image compression model and the second output channel weight of the target fusion output channel corresponding to the target fusion input channel to extract features from the intermediate fused image features to obtain the image fusion features of the sub-band image.
[0087] The fusion input channel layer can be a fusion input corresponding to the features of the intermediate fused image, for example, it could be... Figure 3 The 128 second input channels shown are for each of “Conv1”, “Conv2” and “Conv_post”.
[0088] Multiple second input channels correspond to multiple first output channels. For example, intermediate fusion image features are... Figure 3 The sub-band image features corresponding to the 128 first output channels shown are fused with the reference image features to obtain intermediate fused image features corresponding to the 128 first output channels. Then, the 128 second input channels indicated by the fused input channel layer corresponding to the intermediate fused image features are fused together.
[0089] Multiple second input channels correspond to multiple second output channels. The multiple second input channels include a target fusion input channel, and the multiple second output channels include a target fusion output channel. Each of the multiple second input channels has a different weight, and each of the multiple second output channels has its own weight.
[0090] The number of the second input channel and the second output channel are selected according to actual needs. For example, Figure 3 In the current information extraction network shown, "Conv1" has 128 second input channels and 128 second output channels, and "Conv2" has 128 second input channels and 128 second output channels.
[0091] In the process of processing intermediate fused image features using an image compression model, the intermediate fused image features indicated by multiple second input channels can be processed sequentially. Alternatively, multiple second input channels can be grouped, with each group including at least one second input channel.
[0092] By utilizing the second input channel weight of the target fusion input channel and the second output channel weight of the target fusion output channel corresponding to the intermediate fused image features in the image compression model, feature extraction of intermediate fused image features can be performed. This can be done even when there is only one second input channel and one second output channel. Feature extraction of intermediate fused image features can be performed based on the second input channel weight and the second output channel weight, and the image fusion features corresponding to the second output channel can be output.
[0093] Alternatively, in the case of multiple second input channels and multiple second output channels, for the p-th second input channel and the q-th second output channel out of P second input channels and Q second output channels, feature extraction is performed on the intermediate fused image features and the (p-1)-th image fused features using the weights of the p-th second input channel and the q-th second output channel, to obtain the p-th image fused feature corresponding to the q-th second output channel. This operation is repeated until P image fused features corresponding to each of the Q second output channels are output.
[0094] After determining the multiple image fusion features corresponding to each of the multiple second output channels, for any one of the multiple second output channels, the multiple image fusion features corresponding to that second output channel are fused to determine the image fusion features of the sub-band image by combining the fused image fusion features corresponding to each of the multiple second output channels.
[0095] For the pointwise convolution of "Conv1" and "Conv2", a cross-layer pipeline is used for calculation. For some channels of "Conv1", the results can be directly fed into the channel convolution of another "Conv2", which can further reduce the consumption of external memory access and on-chip resources.
[0096] According to the embodiments of this application, since a specified number of intermediate fused image features corresponding to the second output channels and the second output channels are processed sequentially, the amount of data processed in a single process is smaller than the amount of data processed without splitting the channels, thereby improving the efficiency of processing intermediate fused image features.
[0097] According to an embodiment of this application, feature extraction is performed on the intermediate fused image features using the second input channel weight of the target fused input channel layer corresponding to the intermediate fused image features in the image compression model and the second output channel weight of the target fused output channel corresponding to the target fused input channel, to obtain the image fusion features of the sub-band image. This includes: when there are multiple target fused input channels, performing the following operations: for the m-th target fused input channel among the M target fused input channels, based on the m-th second input channel weight and the m-th second output channel weight, extracting the m-th intermediate fused image channel features and the (m-1)-th image fusion features indicated by the m-th target fused input channel. Feature extraction is performed to obtain the image fusion feature of the m-th image. When m=M, the image fusion feature of the m-th image is determined as the image fusion feature of the sub-band image indicated by multiple target fusion input channels. When m=1, based on the weights of the first second input channel and the weights of the first second output channel, feature extraction is performed on the feature of the first intermediate fusion image channel indicated by the first target fusion input channel to obtain the first image fusion feature. Here, the weights of the M second input channels are different, the weights of the M second output channels are different, 1≤m≤M, and m and M are both integers. Based on the image fusion feature of the sub-band image corresponding to the target fusion input channel of the fusion input channel layer, the image fusion feature of the sub-band image is obtained.
[0098] Figure 5 This diagram illustrates the process of extracting features from intermediate fused image features to obtain sub-band image fusion features according to an embodiment of this application.
[0099] like Figure 5 As shown, with eight second input channels and two second output channels, and Taking the convolution kernel as an example, Ibuf0-Ibuf7 represent intermediate fused image features indicated by different second input channels; W and G indicate two second output channels. For W0 and G0, they respectively indicate the weights of the 0th second input channel and the Wth second output channel, and the weights of the 0th second input channel and the Gth second output channel; for W1 and G1, they respectively indicate the weights of the 1st second input channel and the Wth second output channel, and the weights of the 1st second input channel and the Gth second output channel; (0,0), (0,1), (1,0), and (1,1) correspond to... The convolution kernel indicates the four convolution kernel positions.
[0100] like Figure 5 As shown, W0(0,0) is used to extract features from the intermediate fused image output of Ibuf0 based on the weights of the 0th second input channel, the weights of the Wth second output channel, and the position of the (0,0)th convolutional kernel, to obtain the first convolutional kernel position feature; W0(0,1) is used to extract features from the intermediate fused image output of Ibuf1 and the first convolutional kernel position feature based on the weights of the 0th second input channel, the weights of the Wth second output channel, and the position of the (0,1)th convolutional kernel, to obtain the second convolutional kernel position feature.
[0101] based on Figure 5 The operation shown determines the image fusion features of the sub-band images output by multiple second output channels, fuses the image fusion features of the sub-band images output by multiple second output channels, and determines the image fusion features of the sub-band images.
[0102] According to the embodiments of this application, since a specified number of intermediate fused image features corresponding to the second output channels and the second output channels are processed sequentially, the amount of data processed in a single process is smaller than the amount of data processed without splitting the channels, thereby improving the efficiency of processing intermediate fused image features.
[0103] According to an embodiment of this application, the compression architecture further includes an on-chip storage unit; the method further includes: after obtaining the target image features corresponding to the target fusion input channel, clearing the intermediate data corresponding to the target fusion input channel, and storing the target image features corresponding to the target fusion input channel in the on-chip storage unit, wherein the intermediate data is the data generated during the process of generating the target image features corresponding to the target fusion input channel.
[0104] Intermediate data refers to the data generated during the process of generating target image features corresponding to the target fusion input channel, and may include multiple output results corresponding to the first output channel.
[0105] After obtaining the target image features corresponding to the target fusion input channel, the target image features are stored in the on-chip storage unit, and the intermediate data is cleared. After clearing, the intermediate fusion image features corresponding to the next target fusion input channel are processed until the target image features of each of the multiple target fusion input channels are obtained.
[0106] After obtaining the target image features of each of the multiple target fusion input channels, the target image features of each of the multiple target fusion input channels are extracted from the on-chip storage unit and input into the fusion sub-unit to fuse the target image features of each of the multiple target fusion input channels to obtain the target image features of the sub-band image, and then obtain the image fusion features corresponding to each of the multiple image blocks.
[0107] Figure 6 A schematic diagram illustrating feature extraction of a sub-band image according to an embodiment of this application is shown.
[0108] For example Figure 6 Taking feature extraction of a sub-band image as an example, the sub-band image is processed based on 128 first input channels (first input channel 1, first input channel 2, first input channel 3, ..., first input channel 128). After feature extraction, 128 first output channels (first output channel 1, first output channel 2, first output channel 3, ..., first output channel 128) output 128 sub-band image features respectively. For any one of the 128 first output channels, the 128 sub-band image features are fused to obtain the sub-band image features corresponding to that first output channel, so as to further obtain the sub-band image features of each of the 128 first output channels (sub-band image feature 1, sub-band image feature 2, sub-band image feature 3, ..., sub-band image feature 128). In addition, after obtaining the sub-band image features (sub-band image feature 1, sub-band image feature 2, sub-band image feature 3, ..., sub-band image feature 128) of each of the 128 first output channels, the 128 sub-band image features output by the 128 first output channels are used as intermediate data and cleared.
[0109] In one comparative example of this application, the convolution method is as follows: Y is the output of the network layer, A is the result of multiplying the input data slice with the weights, B is the bias, and C is the scaling factor.
[0110] Assuming that 128 output channels can only calculate 16 input channels at a time, then , This represents the result of multiplying the input data slices from the first 16 input channels with their weights, for the purpose of... , ,but Corresponding to The results are consistent with Consistent. It can be seen that the total computation indicated by the embodiments of this application is consistent with the total computation of the comparative example of this application.
[0111] While keeping the total computing power unchanged, the on-chip storage requirement can be reduced from 8722kb to 3606.75kb, or 41% of the original, by clearing intermediate data.
[0112] According to embodiments of this application, by clearing intermediate data, the on-chip storage space can be reduced, thereby improving the utilization efficiency of on-chip resources and thus enhancing the data processing efficiency of the image compression architecture.
[0113] According to an embodiment of this application, the image to be compressed is divided into multiple sub-band images to obtain multiple image blocks for each sub-band image. This includes: analyzing the bandwidth and storage requirements of a preset initial block area based on the resource data of the image compression architecture, obtaining analysis results, whereby the analysis results characterize whether the resource data of the image compression architecture can provide the bandwidth resources indicated by the bandwidth requirements of the initial block area and the storage resources indicated by the storage requirements of the initial block area. The resource data is the resource data used by the image compression architecture to process the image to be compressed; optimizing the size of the initial block area based on the analysis results to obtain a target block area; and dividing the image to be compressed into multiple sub-band images based on the target block area to obtain multiple image blocks for each sub-band image.
[0114] The method of directly sending the entire image to be compressed into the image compression model for a network layer calculation, writing the result back to DRAM, and then reading it back from DRAM for the next network layer inference is too inefficient, as it involves nearly 70% of the time being spent on data transmission. Therefore, it is necessary to divide the entire image to be compressed into blocks, and then send different image blocks into the image compression model for processing. This internal storage unit can store the intermediate results of the network layers of different image compression models without having to send them to DRAM for reading, thus reducing the time spent accessing external memory.
[0115] Figure 7 A block diagram based on a comparative example of this application is shown; Figure 8 A block diagram according to an embodiment of this application is shown.
[0116] The block-based approach in the scale brings about such Figure 7The overlapping portion between the two image blocks shown (e.g., A1 and A2, A1 and A3) will lead to accuracy loss and incorrect results if the overlapping portion is not considered and the different image blocks (e.g., B1~B4) are directly input into the image compression model for inference. Therefore, when performing convolution on the second image block, the data of the overlapping portion of the first image block needs to be stored on-chip in advance and used together with the second image block for inference to verify the correctness of the result.
[0117] The image to be compressed is divided into multiple sub-bands, and then processed according to the following... Figure 8 The block-based approach shown dynamically divides the feature map into several narrow strips (tiles) along the short side. Unlike proportional block convolution, this approach needs to consider overlap in both row and column directions, increasing the difficulty of implementing network layer fusion. Figure 8 The block segmentation method shown only needs to consider overlap in one direction under the layer fusion data stream. In addition, due to the structural characteristics of the multi-scale pyramid decomposition of the image compression model, the shape of the feature map also changes step by step.
[0118] Because of the adoption Figure 8 The block-based approach shown only requires consideration of row or column overlap during the network layer fusion process. In each convolutional operation, only two columns of overlap data need to be filled at the left boundary of the image block, without considering the filling of the other three boundaries. Considering that the overlap will continuously increase with the number of fusion layers, and that image block consumption will affect the fusion layer depth, we choose to repeatedly load the block data at the input data end (the second image block and the first image block have overlapping parts; taking column-based fusion as an example, the last few columns of the first image block are the first few columns of the second image block).
[0119] The bandwidth and storage requirements of the preset initial block area are analyzed to obtain the results. This analysis can be used to analyze the image compression architecture, the image compression model, and the data flow pattern to determine the computational requirements of the initial block area, the data flow method, and the overlapping regions in convolution. For the image compression model, the focus is on the partial sum of the convolution output channels, and the target block area is calculated based on this information.
[0120] Based on the available storage resources of the image compression architecture, the maximum area budget for a single block is inferred, and several candidate blocks (such as striped or near-block) are generated. Subsequently, the on-chip cache requirements, overlapping storage requirements, and throughput of each candidate block are evaluated.
[0121] Based on the analysis results, the size of the initial block area can be optimized by calculating the input, output, and overlapping data volume of each candidate block to ensure that its memory access requirements do not exceed the bandwidth supported by the image compression architecture.
[0122] Based on the analysis results, optimizing the initial block size can also be achieved by selecting the block partitioning scheme with the shortest computation time and lowest bandwidth requirement, assuming the same computing power, according to computation time and bandwidth requirements. In this case, minimizing overlapping storage becomes the core optimization objective.
[0123] The bandwidth and storage requirements for the target block area are based on Equations (1) to (7).
[0124] (1)
[0125] (2)
[0126] (3)
[0127] (4)
[0128] (5)
[0129] (6)
[0130] (7)
[0131] in, Indicates input memory access, Indicates the height of the target block area. Indicates the width of the target block area. Indicates the number of the first input channel. Indicates output memory access. Indicates the number of the first output channel. Indicates weighted memory access, Indicates the height of the convolution kernel. L represents the width of the convolution kernel, and L represents the number of convolutional layers. This indicates total memory access. Indicates input bandwidth. Indicates the output bandwidth. BW represents the processing time and the total bandwidth.
[0132] According to embodiments of this application, by dividing multiple sub-band images into blocks, the resource data of the image compression architecture can meet the bandwidth and storage requirements of the image blocks, thereby improving the processing efficiency of the image compression architecture while avoiding overload.
[0133] According to an embodiment of this application, the above method further includes: performing discrete wavelet transform on the image to be compressed to obtain images to be compressed with different image resolutions; and using the images to be compressed with different image resolutions as multiple sub-band images of the image to be compressed.
[0134] Figure 9 A schematic diagram of performing discrete wavelet transform on an image to be compressed according to an embodiment of this application is shown; Figure 10 A schematic diagram of a sub-band image according to an embodiment of this application is shown.
[0135] The image compression architecture uses a transform network to perform discrete wavelet transform on the image to be compressed, obtaining images to be compressed at different resolutions. The structure of the transform network is as follows: Figure 9 As shown, The image to be compressed is segmented, and intermediate results are obtained through repeated convolutional structures (multiple lifting structures in the forward transform). After transposition and further segmentation, the final result is obtained. The four largest subband images are obtained, and then this process is repeated for the LL subband to obtain... The four second-largest subband images, and so on, undergo four levels of transformation to obtain, as shown below. Figure 10 The 13 subband images shown (s1LL3, S2HL3, S3LH3, S4HH3, S5HL3, S6LH3, S7HH3, S8HL3, S3LH3, S...) 10 HH3, S 11 HL3, S 12 LH3, S 13 HH3).
[0136] According to embodiments of this application, by performing discrete wavelet transform on the image to be compressed to obtain an image to be compressed with the same resolution, the subsequent feature extraction process can capture image features at different resolution levels, thereby improving the accuracy of feature extraction.
[0137] Figure 11 An image compression architecture according to an embodiment of this application is shown.
[0138] A second aspect of this application provides an image compression architecture, such as Figure 11 As shown, it includes:
[0139] In this embodiment, DRAM represents off-chip memory access unit, Ifm buf represents image buffer to be compressed, Weightbuf represents weight buffer, I buf represents image sub-buffer to be compressed, W buf represents weight sub-buffer, PE represents computing unit, multiple PEs form core computing array, Add buf represents fusion buffer, Psum buf represents storage buffer, Postmodule represents data packet post-processing module, Active buf represents activation buffer, and controller represents controller.
[0140] DRAM is used to store the image to be compressed and to receive the compressed target image.
[0141] The on-chip computing unit performs the following operations: It divides the image to be compressed into multiple sub-band images, resulting in multiple image blocks for each sub-band image. Each sub-band image corresponds to the same block region and has a different image resolution. It then uses the input channel layer corresponding to the image block in the image compression model to perform feature fusion on the multiple sub-band images and a reference image block, obtaining image fusion features corresponding to the image block. The reference image block is determined by fusing multiple target image blocks. Finally, it constructs a target image by fusing the image fusion features corresponding to each of the multiple image blocks. The data size of the target image is smaller than that of the image to be compressed. Finally, it sends the target image to the off-chip access unit.
[0142] Specifically, the DRAM sends the image to be compressed to the Ifm buf, and the Ifm buf divides the image to be compressed into multiple sub-band images, resulting in multiple image blocks for each sub-band image.
[0143] The weight buf is used to obtain the weights of the first input channel, the first output channel, the second input channel, and the second output channel provided by the DRAM.
[0144] The I buf is used to acquire multiple image blocks and input them into the computation subunit PE. The W buf is used to acquire the first input channel weight, the first output channel weight, the second input channel weight, and the second output channel weight and input them into PE. PE is used to perform feature fusion on multiple sub-band images and reference image blocks corresponding to the image blocks using the input channel layer corresponding to the image blocks in the image compression model, so as to obtain the image fusion features corresponding to the image blocks.
[0145] The output of PE is sent to Psum buf and Add buf respectively. Add buf is used to fuse the image fusion features corresponding to multiple image blocks, and Psum buf is used to store the output of PE.
[0146] The Post module is used to construct a target image by fusing the image fusion features corresponding to multiple image blocks.
[0147] Active buf is used for data transfer between off-chip memory access units and on-chip computing units.
[0148] According to embodiments of this application, the weight data required for different arrays (weights of the first input channel, the first output channel, the second input channel, and the second output channel) are loaded into the weight buffer, while the image to be compressed is read into the ifm buffer. The weights of the 16 channels of the current layer are loaded from the weight buffer and written to the corresponding W buffer. Then, feature map data is loaded from the ifm buffer and copied into the I buffer. The portion and result obtained from PE are written to the Psum buffer and the Add buffer. The portion and result in the Add buffer are read and sent to the Post module, where they undergo bias processing, scaling factor multiplication, and nonlinear activation function processing to obtain the final result.
[0149] According to embodiments of this application, by dividing the image to be compressed into multiple sub-band images, the image compression architecture can capture image features at different resolution levels during the compression process. Furthermore, by dividing the multiple sub-band images into blocks, the amount of data processed by the image compression architecture in a single operation is reduced, thus avoiding excessive computational and storage resources. In addition, by utilizing the input channel layer corresponding to the image block in the image compression model to perform feature fusion on the multiple sub-band images and the reference image block, complementary feature information provided by different sub-band images can be combined. Moreover, the target image is constructed on-chip based on the image fusion features corresponding to each of the multiple image blocks, without interaction with off-chip access units. This avoids the problem of low processing efficiency caused by long interaction times, thereby improving processing efficiency while avoiding excessive consumption of computational and storage resources.
[0150] Those skilled in the art will understand that the features described in the various embodiments and / or claims of this application can be combined or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments and / or claims of this application can be combined or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.
[0151] The embodiments of this application have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of this application. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. The scope of this application is defined by the appended claims and their equivalents. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this application, and all such substitutions and modifications should fall within the scope of this application.
Claims
1. An image compression method, characterized in that, Applied to an image compression architecture, the image compression architecture including off-chip memory access units and on-chip computing units, the method includes: The on-chip computing unit receives the image to be compressed provided by the off-chip memory access unit, and performs the following operations using the on-chip computing unit: The image to be compressed is divided into multiple sub-band images, resulting in multiple image blocks for each sub-band image. Each sub-band image corresponds to the same block region and has a different image resolution. The multiple sub-band images are determined based on the following operation: performing a discrete wavelet transform on the image to be compressed to obtain images to be compressed with different image resolutions; and using these images with different image resolutions as the multiple sub-band images of the image to be compressed. For any one of the multiple sub-band images, The input channel layer corresponding to the image block in the image compression model is used to extract features from the sub-band image and the reference image block corresponding to the image block, respectively, to obtain the sub-band image features of the sub-band image and the reference image features of the reference image block. The reference image block is determined based on the following operation: multiple sub-band images corresponding to the image block are sorted in ascending order according to their respective image resolutions. For any sub-band image among the multiple sub-band images, multiple preceding sub-band images located before the sub-band image are determined based on the sorting result, and the multiple preceding sub-band images are fused to obtain the reference image block of the sub-band image. The sub-band image features and the reference image features are fused to obtain intermediate fused image features; Feature extraction is performed on the intermediate fused image features to obtain the image fusion features of the sub-band image; Based on the image fusion features of each of the multiple sub-band images, the image fusion features corresponding to the image blocks are obtained; A target image is constructed by fusing image fusion features corresponding to multiple image blocks, and the data size of the target image is smaller than the data size of the image to be compressed; and The target image is sent to the off-chip memory access unit.
2. The method according to claim 1, characterized in that, The step of extracting features from the sub-band image and the reference image block corresponding to the image block using the input channel layer of the image compression model, to obtain the sub-band image features of the sub-band image and the reference image features of the reference image block, includes: Based on the first input channel weight of the target input channel and the first output channel weight of the target output channel corresponding to the target input channel in the input channel layer, feature extraction is performed on the sub-band image and the reference image block respectively to obtain the sub-band image features of the sub-band image and the reference image features of the reference image block.
3. The method according to claim 2, characterized in that, Feature extraction is performed on the sub-band image to obtain the sub-band image features, including: For the nth convolutional kernel position among the N convolutional kernel positions corresponding to the input channel layer, feature extraction is performed on the sub-band image and the (n-1)th convolutional kernel position features based on the first input channel weight, the first output channel weight, and the nth convolutional kernel position to obtain the nth convolutional kernel position features; When n=N, the position feature of the nth convolutional kernel is determined as the sub-band image feature of the sub-band image. When n=1, the sub-band image is feature extracted based on the first input channel weight, the first output channel weight and the position of the first convolutional kernel to obtain the position feature of the first convolutional kernel. 1≤n≤N, where n and N are both integers.
4. The method according to claim 1, characterized in that, The step of extracting features from the intermediate fused image features to obtain the image fusion features of the sub-band image includes: By utilizing the second input channel weight of the target fusion input channel and the second output channel weight of the target fusion output channel corresponding to the intermediate fused image features in the image compression model, feature extraction is performed on the intermediate fused image features to obtain the image fusion features of the sub-band image.
5. The method according to claim 4, characterized in that, The step of extracting features from the intermediate fused image features using the second input channel weight of the target fused input channel layer corresponding to the intermediate fused image features in the image compression model and the second output channel weight of the target fused output channel corresponding to the target fused input channel, to obtain the image fusion features of the sub-band image, includes: If there are multiple target fusion input channels, the following operations are performed: For the m-th target fusion input channel among M target fusion input channels, based on the weight of the m-th second input channel and the weight of the m-th second output channel, feature extraction is performed on the m-th intermediate fusion image channel feature and the (m-1)-th image fusion feature indicated by the m-th target fusion input channel to obtain the m-th image fusion feature. When m=M, the m-th image fusion feature is determined as the image fusion feature of the sub-band image indicated by multiple target fusion input channels. When m=1, based on the weight of the first second input channel and the weight of the first second output channel, feature extraction is performed on the first intermediate fusion image channel feature indicated by the first target fusion input channel to obtain the first image fusion feature. Here, the weights of the M second input channels are different, the weights of the M second output channels are different, 1≤m≤M, and m and M are both integers. The image fusion features of the sub-band image are obtained based on the image fusion features of the sub-band image corresponding to the target fusion input channel of the fusion input channel layer.
6. The method according to claim 5, characterized in that, The image compression architecture also includes on-chip storage units; The method further includes: After obtaining the target image features corresponding to the target fusion input channel, the intermediate data corresponding to the target fusion input channel is cleared, and the target image features corresponding to the target fusion input channel are stored in the on-chip storage unit. The intermediate data is the data generated during the process of generating the target image features corresponding to the target fusion input channel.
7. The method according to claim 1, characterized in that, The step of dividing the multiple sub-band images of the image to be compressed into blocks, resulting in multiple image blocks for each of the multiple sub-band images, includes: Based on the resource data of the image compression architecture, the bandwidth and storage requirements of the preset initial block area are analyzed to obtain the analysis results. The analysis results characterize whether the resource data of the image compression architecture can provide the bandwidth resources indicated by the bandwidth requirements of the initial block area and the storage resources indicated by the storage requirements of the initial block area. The resource data is the resource data used by the image compression architecture to process the image to be compressed. Based on the analysis results, the size of the initial block area is optimized to obtain the target block area; Based on the target block area, the multiple sub-band images of the image to be compressed are divided into blocks respectively, resulting in multiple image blocks for each of the multiple sub-band images.
8. An image compression architecture, comprising: An off-chip memory access unit is used to store the image to be compressed and to receive the compressed target image; The on-chip computing unit is used to perform the following operations: The image to be compressed is divided into multiple sub-band images, resulting in multiple image blocks for each sub-band image. Each sub-band image corresponds to the same block region and has a different image resolution. The multiple sub-band images are determined based on the following operation: performing a discrete wavelet transform on the image to be compressed to obtain images to be compressed with different image resolutions; and using these images with different image resolutions as the multiple sub-band images of the image to be compressed. For any one of the multiple sub-band images, The input channel layer corresponding to the image block in the image compression model is used to extract features from the sub-band image and the reference image block corresponding to the image block, respectively, to obtain the sub-band image features of the sub-band image and the reference image features of the reference image block. The reference image block is determined based on the following operation: multiple sub-band images corresponding to the image block are sorted in ascending order according to their respective image resolutions. For any sub-band image among the multiple sub-band images, multiple preceding sub-band images located before the sub-band image are determined based on the sorting result, and the multiple preceding sub-band images are fused to obtain the reference image block of the sub-band image. The sub-band image features and the reference image features are fused to obtain intermediate fused image features; Feature extraction is performed on the intermediate fused image features to obtain the image fusion features of the sub-band image; Based on the image fusion features of each of the multiple sub-band images, the image fusion features corresponding to the image blocks are obtained; A target image is constructed by fusing image fusion features corresponding to multiple image blocks, and the data size of the target image is smaller than the data size of the image to be compressed; and The target image is sent to the off-chip memory access unit.
Citation Information
Patent Citations
Image compression method and device, image decompression method and device and image compression and decompression system
CN107801026A
Image compression method and device and image decompression method and device
CN114501011A
Accelerator determination method, image processing method, equipment and storage medium
CN120088345A