Winograd-based image compression quantization method and apparatus
Patent Information
- Application Number
- CN202610911452.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-24
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-06-24
AI Technical Summary
[0005]本申请提供一种基于Winograd的图像压缩量化方法和装置,以解决现有技术使用深度学习框架对卷积量化的粒度过粗,导致效果较差的问题
[0016]由以上技术方案可知,本申请提供一种基于Winograd的图像压缩量化方法,通过将卷积权重变换至Winograd域后按向量级粒度进行量化,同时对激活按通道组粒度进行量化,使量化过程与Winograd卷积加速算法在同一数值域内耦合,避免了空间域量化后再进行Winograd变换所引入的额外失配。在减少乘法复杂度的同时有效控制量化误差,提高了JPEG AI模型的部署效率。
Smart Images

Figure CN122453950B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image compression technology, and in particular to an image compression quantization method and apparatus based on Winograd. Background Technology
[0002] With the continuous growth in demand for image content storage and dissemination, minimizing bitrate while ensuring subjective image quality has become a core requirement in the field of image coding. To address this need, the Joint Photographic Experts Group (JPEG) standardization committee launched the JPEG AI project to develop image coding specifications. JPEG AI uses neural networks as the core to build image compression technology.
[0003] The JPEG AI compression framework uses an analysis transform network to map the input image into latent variables, and then generates a bitstream through hyper-prior modeling, context modeling, and entropy coding. At the decoding end, the image is recovered through inverse quantization, hyper-prior decoding, and a synthesis transform network. In this encoding and decoding process, convolution and transposed convolution are the most important operators. Especially in high-resolution image scenarios, the large number of convolutional layers in the backbone and hyper-prior networks leads to high computational cost and memory access pressure.
[0004] To improve computational speed, existing technologies utilize the built-in quantization functions of deep learning frameworks to quantize the weights and activations of convolutions layer by layer in the spatial domain, thereby increasing computational speed by reducing computational cost. However, the quantization granularity of the deep learning frameworks used in existing technologies is too coarse, thus limiting their application to general visual tasks such as classification or detection, resulting in poor quantization performance and unsuitability for image compression tasks. Summary of the Invention
[0005] This application provides an image compression and quantization method and apparatus based on Winograd to solve the problem that the granularity of convolutional quantization using deep learning frameworks in the prior art is too coarse, resulting in poor performance.
[0006] The first aspect of this application provides an image compression and quantization method based on Winograd, the method comprising: Obtain the input parameters of the initial convolution of the target neural network model; the input parameters include activation data and weight data in the original spatial domain; the initial convolution is used to compress and encode the image processed by the target neural network model; Obtain the kernel parameters of the initial convolution; the kernel parameters include the kernel stride; Based on the convolution kernel stride, the weight data is transformed from the original spatial domain to the Winograd domain to obtain a weight tensor in the Winograd domain; wherein, when the convolution kernel stride is 1, a first Winograd weight transformation is performed on the weight data to obtain a weight tensor in a first format; when the convolution kernel stride is 2, a second Winograd weight transformation is performed on the weight data to obtain a weight tensor in a second format. Calculate the weight scale of the weight tensor in the Winograd domain, and quantize the weight tensor according to the weight scale to obtain the quantized weights in the Winograd domain. Calculate the activation scale corresponding to the activation data, and quantize the activation data according to the activation scale to obtain quantized activation; The weight scale, quantization weight, activation scale, and quantization activation in the Winograd domain are input into an integer convolution kernel in the Winograd domain, and an integer convolution operation is performed on the integer convolution kernel to obtain the convolution operation result. Based on the weight scale and the activation scale, the convolution operation result is reverse-scaled to obtain restored data in integer format; The restored data is then superimposed with bias terms to obtain compressed image features in integer format.
[0007] In some feasible embodiments, obtaining the input parameters of the initial convolution of the target neural network model includes: Determine the initial convolutions to be quantized in the target neural network model; the initial convolutions include ordinary convolutions and transposed convolutions, with ordinary convolutions being convolutions that do not require transposition; Obtain the activation tensor data and convolution kernel weight data corresponding to the initial convolution in the original spatial domain; the activation tensor data and convolution kernel weight data are in floating-point format.
[0008] In some feasible embodiments, when the convolution kernel stride is 1, a first Winograd weight transformation is performed on the weight data, including: With the convolution kernel stride being 1, each column of the weight data is processed based on the left multiplication transformation function to map each column of data into a four-dimensional intermediate vector; The weighted data is processed based on the right multiplication transformation function to map each row of data into a four-dimensional output vector. A four-row, four-column matrix is generated based on the four-dimensional intermediate vector and the four-dimensional output vector; the four-row, four-column matrix contains a four-dimensional weight tensor; the four dimensions include the output channel dimension, the input channel group dimension, the Winograd transform grid width dimension, and the Winograd transform grid height dimension.
[0009] In some feasible embodiments, performing a second Winograd weight transformation on the weight data when the convolution kernel stride is 2 includes: When the convolution kernel stride is 2, the convolution kernel corresponding to the weight data is detected; If the convolution kernel is a transposed convolution and the kernel size is 4×4, then the 4×4 weight data is divided into four 2×2 sub-weight data; if the convolution kernel is not a transposed convolution, or the kernel size is not 4×4, then the weight data is divided into four sub-weight data of sizes: 2×2, 2×1, 1×2, and 1×1. The sub-weight data is processed using a left-multiplication transformation function to map each column of data into a five-dimensional intermediate vector; and the sub-weight data is processed using a right-multiplication transformation function to map each row of data into a five-dimensional output vector. A five-dimensional weight tensor is generated based on the five-dimensional intermediate vector and the five-dimensional output vector; the second format includes the output channel dimension, the input channel group dimension, the Winograd transform grid width dimension, the Winograd transform grid height dimension, and the branch dimension.
[0010] In some feasible embodiments, the weight scale of the weight tensor in the Winograd domain is calculated, including: Determine the multiple dimensions of the weight tensor in the Winograd domain; The weight tensor is divided into multiple local weight vectors based on information from multiple dimensions; Calculate the maximum absolute value of each of the multiple local weight vectors, and set the maximum absolute value as the weight scale of the corresponding local weight vector.
[0011] In some feasible embodiments, the weight tensor is quantized according to a weight scale, including: Based on the weight scale corresponding to the local weight vector, the local weight vector is scaled. Based on the preset bit width, the scaled local weight vector is sequentially rounded and clipped to obtain the quantized weights corresponding to the local weight vector.
[0012] In some feasible embodiments, the activation scale corresponding to the activation data is calculated, including: Obtain the input channel information corresponding to the activated tensor data; Based on the input channel information, the activation tensor data is divided into multiple input channel groups to obtain multiple grouped activation tensors; Obtain the maximum absolute value of the activation tensor data corresponding to multiple input channels within each group activation tensor, and set the maximum absolute value as the activation scale corresponding to each group activation tensor.
[0013] In some feasible embodiments, the activation data is quantified according to an activation scale, including: Based on the activation scale corresponding to each group activation tensor, the activation tensor data corresponding to each group activation tensor is scaled. Based on the preset bit width, the scaled activation tensor data is sequentially rounded and pruned to obtain the quantized activation corresponding to the grouped activation tensor.
[0014] In some feasible embodiments, integer convolution operations are performed on the integer convolution kernel, including: The Winograd algorithm is used to perform integer convolution operations on the quantized weights and quantized activations in the integer convolution kernel.
[0015] Secondly, this application provides an image compression and quantization apparatus based on Winograd, the apparatus comprising: The parameter acquisition module is configured to acquire the input parameters of the initial convolution of the target neural network model; the input parameters include activation data and weight data in the original spatial domain; the initial convolution is used to compress and encode the image processed by the target neural network model; The kernel parameter acquisition module is configured to acquire the kernel parameters of the initial convolution; the kernel parameters include the kernel stride. The weight transformation module is configured to transform the weight data from the original spatial domain to the Winograd domain based on the convolution kernel stride, to obtain a weight tensor in the Winograd domain; wherein, when the convolution kernel stride is 1, a first Winograd weight transformation is performed on the weight data to obtain a weight tensor in a first format; when the convolution kernel stride is 2, a second Winograd weight transformation is performed on the weight data to obtain a weight tensor in a second format. The weight quantization module is configured to calculate the weight scale of the weight tensor in the Winograd domain, and to quantize the weight tensor according to the weight scale to obtain the quantized weights in the Winograd domain. The activation quantization module is configured to calculate the activation scale corresponding to the activation data, and to quantize the activation data according to the activation scale to obtain quantized activation. The integer convolution operation module is configured to input the weight scale, the quantization weight, the activation scale, and the quantization activation in the Winograd domain into the integer convolution kernel in the Winograd domain, and perform integer convolution operation on the integer convolution kernel to obtain the convolution operation result. The restoration module is configured to reverse scale the convolution operation result based on the weight scale and the activation scale to obtain restored data in integer format; The bias term overlay module is configured to overlay bias terms on the restored data to obtain compressed image features in integer format.
[0016] As can be seen from the above technical solutions, this application provides an image compression and quantization method based on Winograd. By transforming the convolution weights to the Winograd domain and then quantizing them at the vector level, while simultaneously quantizing activations at the channel group level, the quantization process is coupled with the Winograd convolution acceleration algorithm within the same numerical domain. This avoids the additional mismatch introduced by performing Winograd transformation after spatial domain quantization. While reducing multiplication complexity, it effectively controls quantization errors and improves the deployment efficiency of the JPEG AI model. Attached Figure Description
[0017] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A flowchart illustrating the image compression and quantization method based on Winograd in some embodiments is shown; Figure 2 A schematic diagram of the structure of an image compression and quantization device based on Winograd is shown in some embodiments. Detailed Implementation
[0019] The technical solutions of the embodiments of this application will now be described with reference to the accompanying drawings. To facilitate a clear description of the technical solutions of the embodiments of this application, the use of terms such as "first," "second," etc., in the embodiments of this application is for illustrative purposes and to distinguish the objects being described. There is no particular order between them, nor does it indicate a specific limitation on the number of devices in the embodiments of this application, and they do not constitute any limitation on the embodiments of this application.
[0020] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of this application.
[0021] It should be noted that many specific details are set forth in the following description in order to provide a full understanding of this application. However, this application may also be implemented in other ways different from those described herein. Therefore, the scope of protection of this application is not limited to the specific embodiments disclosed below.
[0022] The following explanations of the technical terms mentioned in the embodiments of this application are provided to facilitate understanding by those skilled in the art.
[0023] JPEG AI is an image coding standardization project initiated by the Joint Photographic Experts Group (JPEG) standardization committee to develop image coding specifications. Its core idea is to replace the Discrete Cosine Transform (DCT) and quantization modules in traditional codecs with neural networks, aiming to reduce bitrate while maintaining subjective image quality, and providing a representation suitable for direct machine vision input data. The JPEG AI standard uses analytical transform networks and synthetic transform networks as core components of the encoder and decoder, capturing the statistical dependencies of latent variables through advanced prior modeling, and then combining autoregressive context models and entropy coding to achieve efficient compression.
[0024] Quantization refers to the process of mapping floating-point values to discrete values with finite precision. In neural network inference deployment, quantization typically involves converting 32-bit floating-point (float32) weights and activation values into low-bit integers (such as int8 or int4) to reduce the model's storage overhead and computational cost. The core parameter in the quantization process is the scaling factor, which determines the mapping relationship between floating-point values and integer values. The finer the granularity of the scaling factor, the closer the distribution of the quantized values is to the original floating-point distribution, and the smaller the quantization error.
[0025] In neural networks, activation refers to the intermediate feature maps generated by the operations of each layer. In convolutional operations, the output feature map of the previous layer becomes the input activation of the current convolutional layer.
[0026] In neural networks, weights refer to the numerical values at each position in the convolution kernel (filter). They determine how much "importance" is assigned to pixels (or features) at different positions in the input local region during sliding convolution, thereby extracting specific features.
[0027] Weight quantization is the process of mapping floating-point format convolution kernel weight data to low-bit integer representation through scaling, rounding, and pruning operations. The purpose is to reduce the amount of weight storage and support faster integer operations.
[0028] Activation quantization: The process of quantizing the input or output feature maps (activation data) of each layer in the neural network inference process from floating-point format to integer format.
[0029] Scale: A scaling factor used in quantization operations to map floating-point numerical fields to integer numerical fields. Its value directly determines the resolution of the quantization.
[0030] Transposed convolution, also known as deconvolution, is a type of convolution operation used for upsampling. It is widely used in the synthesis transform network at the image compression and decoding end. Its spatial size expansion effect is opposite to the downsampling direction of ordinary convolution.
[0031] Inverse scaling: The operation of multiplying the integer convolution output by the product of the activation scale and the weight scale to restore the feature map from the integer domain to the floating-point format.
[0032] The Winograd algorithm is a mathematical algorithm for optimizing convolution operations, aiming to reduce computation and accelerate the convolution process. This algorithm decomposes the standard convolution into a series of smaller sub-convolutional blocks and employs a computational flow of pre-transformation, element-wise multiplication, and inverse transform, reducing the number of multiplication operations required for convolution. For example, a one-dimensional convolution F(2,3) requires six multiplications in ordinary matrix multiplication, while the Winograd algorithm achieves one-dimensional convolution with only four multiplications, resulting in a 1.5x speedup.
[0033] The application scenarios of this application will be explained below.
[0034] With the continuous growth in demand for image content storage and dissemination, minimizing bitrate while ensuring subjective image quality has become a core requirement in the field of image coding. To address this need, the Joint Photographic Experts Group (JPEG) standardization committee launched the JPEG AI project to develop image coding specifications.
[0035] The JPEG AI compression framework uses an analysis transform network to map the input image into latent variables, and then generates a bitstream through hyper-prior modeling, context modeling, and entropy coding. At the decoding end, the image is recovered through inverse quantization, hyper-prior decoding, and a synthesis transform network. In this encoding and decoding process, convolution and transposed convolution are the most important operators. Especially in high-resolution image scenarios, the large number of convolutional layers in the backbone and hyper-prior networks leads to high computational cost and memory access pressure.
[0036] Currently, the official JPEG AI standard has not provided a quantized version of the network for practical inference deployment. In the absence of an official standard quantized network, existing technologies typically improve computational speed by relying on the built-in quantization functions of deep learning frameworks to perform conventional layer-by-layer or channel-by-channel quantization on convolutional weights and activations in the spatial domain, or by directly attempting integer quantization strategies.
[0037] However, conventional neural network quantization schemes often face serious technical bottlenecks in image compression tasks. Existing deep learning frameworks, typically PyTorch, suffer from issues such as excessively coarse quantization granularity and mismatch between data layout and hardware execution. Therefore, these quantization schemes are generally used for general classification or detection tasks and cannot be adjusted for the unique structural features of image compression networks. They cannot achieve fast computation while maintaining accuracy, and are prone to prediction bias beyond the prior scale, leading to probabilistic model mismatch, increased performance fluctuations between different bitrates, and increased visual distortion or striping artifacts on the decoding side. Therefore, their quantization performance is poor and unsuitable for image compression tasks.
[0038] To address the aforementioned issues, this application provides an image compression and quantization method based on Winograd. This method transforms convolution weights to the Winograd domain and then quantizes them at the vector level, while simultaneously quantizing activations at the channel group level. This couples the quantization process with the Winograd convolution acceleration algorithm within the same numerical domain, avoiding the additional mismatch introduced by performing Winograd transformation after spatial domain quantization. By reducing multiplication complexity and effectively controlling quantization errors, this method improves the deployment efficiency of the JPEG AI model.
[0039] It's important to note that the quality of quantization directly determines the compression performance. The metric for compression performance is "rate-distortion"—lower bit rate (R) and lower distortion (D) are better. A larger quantization step size results in fewer bits but higher distortion, while a smaller quantization step size results in lower distortion but more bits. The optimal quantization strategy minimizes distortion given a bit budget, or minimizes bits given a distortion constraint. This invention is a quantization strategy specifically designed for rate-distortion optimization in image compression.
[0040] The essence of the Winograd transform is to convert spatial domain convolution into element-wise multiplication in the transform domain, thereby reducing the number of multiplications. Traditional methods quantize in the pixel domain, where the quantization error is uniform across each pixel, failing to distinguish the different impacts of different locations on the output. The Winograd transform maps the convolution kernel to a larger transform domain space. Each element in the transform domain is a specific linear combination of the original convolution kernel parameters, and different elements contribute to the final convolution output in different ways—combinations at certain locations have a large impact on the output, while those at other locations have a small impact. If quantization is performed in the transform domain, different quantization precisions can be assigned to these combinations—combinations with a large impact on the output retain more precision, while combinations with a small impact are quantized more aggressively.
[0041] Regarding convolution output, the quantization results of this invention have two significant features that distinguish them from traditional quantization schemes.
[0042] One is spatial block-level correlation. Due to the linear mixing effect of the Winograd inverse transform, the quantization error sources in the transform domain are redistributed to multiple pixels in the output block by the inverse transform matrix, so that each pixel in the same output block shares the same error source, and the quantization errors are no longer spatially independent - this is a structural feature that traditional pixel-domain quantization (independent and identically distributed errors) does not have.
[0043] Secondly, there is a non-uniform error distribution. Because elements at different positions in the transform domain are quantized differently, elements that have a greater impact on the output have higher precision, while elements that have a smaller impact have lower precision. As a result, the quantization error is distributed non-uniformly among the pixels in the output block.
[0044] Figure 1 The diagram illustrates the flow of an image compression and quantization method based on Winograd in some embodiments. For example... Figure 1 As shown, the image compression and quantization method based on Winograd provided in this application includes the following steps S100 to S800.
[0045] Step S100: Obtain the input parameters of the initial convolution of the target neural network model. The input parameters include activation data and weight data in the original spatial domain; the initial convolution is used to compress and encode the image processed by the target neural network model.
[0046] In this embodiment, the target neural network model is a JPEG AI neural image compression model. This model includes an analysis transform network and a synthesis transform network, containing a large number of convolutional layers, including ordinary convolutions and transposed convolutions. Ordinary convolutions are those that do not require transposition. When quantizing the JPEG AI model, the convolutions to be quantized can be determined first; in this embodiment, these are referred to as the initial convolutions. Subsequent Winograd transformations and quantization processes can be performed on the initial convolutions to reduce the computational load.
[0047] For each initial convolution, its corresponding input parameters can be obtained. In neural networks, the input parameters of a convolution refer to its activation data and the weight data corresponding to the convolution kernel in the spatial domain. Both data are in floating-point format, such as float32. The activation data can be determined by the input feature map, and the convolution kernel weights can be determined by the convolution kernel parameters.
[0048] Taking a typical JPEG AI backbone convolutional layer as an example, its input activation data can include the number of input channels, the number of output channels, height, and width. The weight data of the convolutional kernel can include the number of output channels, the number of input channels, the kernel height, the kernel width, and the stride.
[0049] Step S200: Obtain the kernel parameters for the initial convolution. The kernel parameters include the kernel stride.
[0050] Step S300: Convert the weight data in the original spatial domain into a weight tensor in the Winograd domain. This includes: based on the convolution kernel stride, converting the weight data from the original spatial domain to the Winograd domain to obtain a weight tensor in the Winograd domain; wherein, when the convolution kernel stride is 1, performing a first Winograd weight transformation on the weight data to obtain a weight tensor in a first format; and when the convolution kernel stride is 2, performing a second Winograd weight transformation on the weight data to obtain a weight tensor in a second format.
[0051] Specifically, the weight data is the floating-point weight tensor of the neural network data in the spatial domain. The weight data of each initial convolution is transformed by Winograd weight transformation and mapped to the Winograd domain, so as to lay the foundation for efficient quantization by utilizing the characteristics of the Winograd domain.
[0052] The initial convolution parameters, including the convolution stride, can be obtained. Depending on the convolution stride, different Winograd weight transformation formats are performed on the weight data.
[0053] Through the above transformation method, the spatial domain convolution kernel completes the intra-domain rearrangement consistent with the Winograd computation structure before low-bit computation, ensuring that convolutional layers in different time length modes can adapt to the subsequent Winograd domain quantization process.
[0054] Step S400: Calculate the weight scale of the weight tensor, and quantize the weight tensor according to the weight scale to obtain the quantized weight.
[0055] For a weight tensor in the Winograd domain, we can first determine various dimensional information corresponding to the weight tensor. For example, these dimensional information include the output channels, the input channel group, and the position within the Winograd transform grid. When the convolution stride is 2, the dimensional information also includes the stride branch index.
[0056] Based on the aforementioned multi-dimensional information, the local weight vector partitioning method can be determined, thereby dividing the weight tensor into multiple local weight vectors. For example, each local weight vector can correspond to a specific combination of output channel, input channel group, and Winograd transformation grid position (and possible step size branches).
[0057] For each local weight vector, its corresponding scale can be calculated independently, referred to as the weight scale in this embodiment. The scale represents the scaling factor that maps the floating-point numerical field to the integer numerical field, and thus can characterize the dynamic range of the vector during the quantization process. The smaller the scale value, the higher the accuracy after quantization; the larger the scale value, the lower the accuracy after quantization.
[0058] By using the weight scale of each local weight vector, the local weight vector is scaled and quantized to obtain the quantized weight data corresponding to each local weight vector.
[0059] By using the above vector-wise scaling method, the system can more accurately reflect the dynamic range characteristics of each local position of the Winograd domain weights than traditional layer-wise or channel-wise quantization, effectively reducing quantization errors.
[0060] Step S500: Calculate the activation scale corresponding to the activation data, and quantize the activation data according to the activation scale to obtain quantized activation.
[0061] In this embodiment, the input channel information corresponding to the activation data can be obtained, and the activation data can be divided into multiple input channel groups along the channel dimension based on the input channel information to obtain multiple grouped activation tensors. Each group is processed as a unit. For example, every 16 input channels constitute one channel group.
[0062] The scaling of activation data also follows the Winograd domain adaptation principle. Based on the input channel group partitioning strategy, a scaling factor corresponding to each activation group is determined, referred to as the activation scaling in this embodiment. That is, each channel group has an independent activation scaling factor. The grouping granularity can be strictly aligned with the channel groups in weight quantization, ensuring that activation and weights are scaled in the Winograd domain, thereby guaranteeing the consistency and predictability of the numerical range in subsequent integer multiplication and addition operations.
[0063] After calculating the activation scale, the activation data is quantized based on the activation scale. The activation data for each group can be scaled according to the activation scale to obtain the quantized activation data for each group.
[0064] By scaling and quantizing activations by channel group, the problem of partial channel saturation or insufficient quantization accuracy caused by large differences in the numerical distribution of different input channels is effectively avoided. At the same time, the data layout of the quantized activation tensor is consistent with the input requirements of the Winograd integer convolution kernel, allowing it to directly participate in integer convolution operations within the Winograd domain.
[0065] Step S600: Input the weight scale, the quantization weight, the activation scale, and the quantization activation in the Winograd domain into an integer convolution kernel in the Winograd domain, and perform integer convolution operation on the integer convolution kernel to obtain the convolution operation result.
[0066] After activation quantization and weight quantization are completed, the quantized activation, quantized weight, and corresponding weight scale and activation scale are input into a predefined Winograd integer convolution kernel, which is also referred to as a computation kernel in this embodiment.
[0067] During execution, the quantized data can be padded and divided into blocks according to the convolution parameters, and then integer multiplication and addition operations can be performed on the integer convolution kernel in the Winograd transform domain. The Winograd algorithm can be used to perform integer convolution operations on the quantized weights and quantized activations in the integer convolution kernel. Specifically, the convolution operation that requires 3×3 multiplications in the spatial domain is transformed into an element-wise multiplication and addition operation in the Winograd domain that requires only a few multiplications.
[0068] Since the entire computation path performs integer multiplication and addition operations directly on the Winograd field, it avoids the additional numerical mismatch introduced by performing Winograd transformation after spatial field quantization. Therefore, it can maintain more stable numerical performance in low bit-width deployments, effectively reducing the amount of computation and improving the computation speed.
[0069] Step S700: Based on the weight scale and activation scale, the convolution operation result is reverse scaled to obtain the restored data in integer format.
[0070] The system performs reverse scaling on the integer convolution results in the Winograd domain based on weight and activation scales, restoring the integer results to integer format data.
[0071] Step S800: Add bias terms to the restored data to obtain compressed image features in integer format.
[0072] After the data is restored and a bias term is added, it becomes the compressed image features corresponding to the target neural network model. The restored compressed image features can be directly input into subsequent network layers for inference, and can be aligned with the accuracy of the original model without additional calibration, significantly reducing the fine-tuning cost during deployment.
[0073] Through the above process, the output feature map of the convolutional layer after quantization replacement is numerically consistent with the output of the original network's floating-point convolutional layer, ensuring that the numerical propagation of subsequent layers is not affected by quantization replacement.
[0074] Based on the above technical solution, quantization processing and the Winograd acceleration algorithm can be coupled in the same numerical domain, reducing errors while ensuring that weight and activation values are compressed to lower bits. This method has good compatibility with existing convolutional networks, and can replace convolutional layers with Winograd quantized convolutional layers without changing the original network topology, thereby non-intrusively improving convolutional inference efficiency and enhancing low-bit deployment performance.
[0075] The above is the overall flowchart of the image compression and quantization method based on Winograd in this application. The specific implementation methods of the above steps are explained in detail below with reference to the accompanying figures.
[0076] In step S100, the input parameters of the initial convolution of the target neural network model are obtained.
[0077] In the JPEG AI neural image compression network, convolution operations are distributed across multiple layers of the analysis transform network, the synthesis transform network, and the advanced prior network. Based on the functional implementation path, this method categorizes the convolution operators in the network into three types: (1) ordinary convolution, i.e., standard convolution with a stride of 1 or 2, responsible for feature extraction and downsampling; (2) transposed convolution, i.e., transposed (inverse) convolution with a stride of 1 or 2, responsible for feature upsampling, widely used in the synthesis transform network; (3) irregular convolution kernels, convolution kernels with shapes other than standard squares (e.g., strip convolution kernels, dilated convolution kernels, and deformable convolution kernels), whose size is usually not a square structure such as 1×1 or 3×3. In the embodiments of this application, such convolution kernels are defined as not suitable for the Winograd transform acceleration algorithm and require maintaining the original floating-point quantization path of the convolution operation. Therefore, irregular convolution kernels are not included in the Winograd transform process, maintaining the original floating-point quantization path of the JPEG AI neural image compression network. In some embodiments, the irregular convolutional kernel can be the first and last convolutional layers of the entire network. The irregular convolutional kernel can be embodied in the convolutional process that performs concatenation (i.e., channel expansion) operation in the input first layer.
[0078] In other words, the embodiments of this application uniformly incorporate the first two types of convolution, namely ordinary convolution and transposed convolution, into the Winograd integer calculation process and set them as the initial convolution, that is, the convolution to be quantized.
[0079] In the target neural network model, initial convolutions that satisfy the Winograd quantization condition can be identified (e.g., convolution kernel size of 3×3, stride of 1 or 2, ordinary convolution or transposed convolution). For each identified convolutional layer to be quantized, the system extracts its convolution kernel weight data in floating-point format in the spatial domain as weight data, and obtains the input activation tensor data as activation data; both are floating-point format data.
[0080] The above classification and recognition mechanism can accurately select convolutional layers suitable for the Winograd transform. The convolution direction identifier `transposed` can also distinguish between ordinary convolutions and transposed convolutions, providing correct input parameters for subsequent Winograd transform steps. Specifically, when `transposed` is True, the transposed convolution path is triggered; when `transposed` is False, the ordinary convolution path is executed.
[0081] In step S200, the kernel parameters for the initial convolution are obtained. These kernel parameters include the kernel stride.
[0082] The system reads the kernel parameters of the current convolutional layer to be quantized, including the kernel stride, the transposed orientation (transposed), padding, and the number of groups. A unified weight transformation entry function is set up, automatically distinguishing the transformation paths of ordinary convolutions and transposed convolutions based on the transposed flag.
[0083] In step S300, the weight data in the original spatial domain is transformed using the Winograd weight transformation to convert it into a weight tensor in the Winograd domain. Specifically, based on the convolution kernel stride, the weight data is transformed from the original spatial domain to the Winograd domain to obtain a weight tensor in the Winograd domain.
[0084] When performing Winograd weight transformation, different data formats can be generated for different step sizes.
[0085] With a kernel stride of 1, a first Winograd weight transformation is performed on the weight data to obtain a weight tensor in a first format. This first format is four-dimensional, including the output channel dimension, the input channel group dimension, the Winograd transformation grid width dimension, and the Winograd transformation grid height dimension.
[0086] Step S300 specifically includes the following steps: With the convolution kernel stride being 1, each column of the weight data is processed based on the left multiplication transformation function to map each column of data into a four-dimensional intermediate vector; The weighted data is processed based on the right multiplication transformation function to map each row of data into a four-dimensional output vector. A four-row, four-column matrix is generated based on the four-dimensional intermediate vector and the four-dimensional output vector; the four-row, four-column matrix contains a four-dimensional weight tensor.
[0087] In some embodiments, when the stride parameter is set to 1, the Winograd weight transformation is performed on the two-dimensional slice of the input convolution kernel. This transformation process includes two stages: column transformation and row transformation. In the column transformation stage, each column of the kernel slice is processed by iteratively calling the left-multiply transformation function (WinoWgtTransformLeft), mapping the data of each column to a four-dimensional intermediate vector, and writing it into a four-row, four-column intermediate storage matrix according to the column index. The effective transformation result of this matrix is only distributed in the first two columns, while the last two columns are set to zero to unify the dimension. In the row transformation stage, each row of this intermediate matrix is processed iteratively, and the data of each row is mapped to a four-dimensional output vector by calling the right-multiply transformation function (WinoWgtTransformRight), and written into the final four-row, four-column coefficient matrix according to the row index, thus obtaining a total of sixteen transformation coefficients. After completing the bidirectional transformation, all sixteen coefficients are extracted from the coefficient matrix in row-major order. Based on the channel storage priority indicated by the preset transpose layout flag (i.e., output channel priority or input channel priority), the coefficients are sequentially written to the storage offset position corresponding to the output tensor for subsequent Winograd convolution operations.
[0088] It should be noted that the dimensions of the Winograd transform grid depend on the Winograd transform form used. Taking the F(2,2) transform as an example, the convolution operation of a 3×3 convolution kernel with a 2×2 input block is decomposed into 4×4 transform domain elements, so the width and height dimensions of the Winograd transform grid are both 4.
[0089] For example, for a 3×3 convolution with stride=1 (including ordinary convolution and transposed convolution), a 4×3 transformation matrix G is used to perform a two-dimensional transformation on the spatial domain convolution kernel weights W, one for each output channel and one for each input channel group, to obtain the Winograd domain weight tensor of the above four dimensions. Compared with the original weights, the amount of data increases after the transformation, but the dynamic range of each transformed grid position is more focused, which is beneficial for quantization accuracy control.
[0090] With a kernel stride of 2, the Winograd weight transformation is performed on the weight data to obtain a weight tensor in a second format. This second format is five-dimensional, including the output channel dimension, the input channel group dimension, the Winograd transformation grid width dimension, the Winograd transformation grid height dimension, and the branch dimension.
[0091] The branch dimension corresponds to the even / odd sampling paths when the stride is 2. Since there are 4 different spatial sampling combinations in the input feature map when the stride is 2, such as odd rows and odd columns, odd rows and even columns, even rows and odd columns, and even rows and even columns, an additional branch dimension is needed to distinguish these 4 sampling branches, so that each branch has an independent weight tensor in the Winograd domain.
[0092] In other words, when the stride is 1, a four-dimensional Winograd weight tensor is output, and when the stride is 2, a five-dimensional tensor with branch dimensions is output, so that the spatial domain convolution kernel completes the intra-domain rearrangement consistent with the Winograd computation structure before low-bit computation.
[0093] With the convolution kernel stride being 2, a second Winograd weight transformation is performed on the weight data, including the following steps: Detect the convolution kernel corresponding to the weight data; If the convolution kernel is a transposed convolution and the kernel size is 4×4, then the 4×4 weight data is divided into four 2×2 sub-weight data; if the convolution kernel is not a transposed convolution, or the kernel size is not 4×4, then the weight data is divided into four sub-weight data of sizes: 2×2, 2×1, 1×2, and 1×1. The sub-weight data is processed using a left-multiplication transformation function to map each column of data into a five-dimensional intermediate vector; and the sub-weight data is processed using a right-multiplication transformation function to map each row of data into a five-dimensional output vector. A five-dimensional weight tensor is generated based on the five-dimensional intermediate vector and the five-dimensional output vector; the second format includes the output channel dimension, the input channel group dimension, the Winograd transform grid width dimension, the Winograd transform grid height dimension, and the branch dimension.
[0094] In some embodiments, the process of converting the convolution kernel weight tensor into the Winograd domain, based on specific parameter analysis, can be represented by the following steps: First, the input tensor is reorganized in contiguous memory and its four-dimensional structure (output channel number oc, input channel number ic, kernel height kh, kernel width kw) is verified. Then, the input channel number is padded to a multiple of 32 (tiled_ic), and the total number of computation blocks is calculated as ic_tiles × oc, with each block containing 32 threads. Based on the stride parameter, if stride == 1, an output tensor of shape [oc, ic, 4, 4] is allocated; if stride = 2, an output tensor of shape [oc, ic, 4, 4, 4] is allocated. Then, the kernel function WinowtTransformKernel is started.
[0095] Within the kernel function, each thread block processes a pair of output channel indices oc_idx and input channel group indices ic_idx. First, based on the transposed flag, kh×kw weight elements are loaded from the input tensor into the register array tmp_in according to the corresponding layout. If stride == 1, WinowgtTransformLeft (left multiplication by the transformation matrix) is called on each column of the kernel to obtain an intermediate matrix, and then WinowgtTransformRight (right multiplication by the transformation matrix) is called on each row to obtain 16 transformation coefficients, which are then sequentially written to the output tensor. If stride == 2, further distinctions are made: when it is a transposed convolution and the kernel size is 4×4, the 4×4 weights are divided into four 2×2 sub-blocks, which are then transformed left and right and written to different offset regions of the output tensor; otherwise, they are processed sequentially according to four sub-kernel shapes: 2×2, 2×1, 1×2, and 1×1. The corresponding elements are extracted from tmp_in (padded with zeros if necessary), a one-dimensional transformation for the sub-kernel size is called, and the 16 transformation coefficients for each round are written to predefined offset segments (0, 16, 32, 48). The above left and right transformation functions all extract a four-element vector and call WinowtTransform1d to implement a one-dimensional Winograd transformation, ultimately mapping the convolution kernel to the Winograd domain to support the fast calculation of subsequent convolutions with different lengths and transposes.
[0096] In step S400, the weight scale of the weight tensor in the Winograd domain is first calculated. The weights are divided into local vectors in the Winograd domain, and the scale is calculated vector by vector.
[0097] After completing the Winograd weight transformation, significant differences exist in the dynamic range of values at different output channels, different input channel groups, and different transformation grid positions within the transform domain weight tensor. If a single scale (layer-by-layer quantization) is applied to the entire weight tensor, local weights with small dynamic ranges will suffer from insufficient precision due to excessively large truncation thresholds. Conversely, if the scale is divided only according to the output channels (channel-by-channel quantization), the differences in dynamic range at various positions within the Winograd transformation grid will be imperceptible, similarly introducing unnecessary quantization errors.
[0098] Therefore, this method employs a structured vector-level scaling design. Specifically, it includes the following steps: Determine the multiple dimensions of the weight tensor. Different tensor lengths correspond to different numbers of dimensions, allowing you to determine the dimensional information for each weight tensor.
[0099] The weight tensor is divided into multiple local weight vectors based on information from multiple dimensions. Each dimension can serve as an index, which is used as the basis for vector partitioning.
[0100] In some embodiments, four dimensions are used as examples: output channel dimension, input channel group dimension, Winograd transform grid width dimension, and Winograd transform grid height dimension. The information quantity of the four dimensions can be multiplied to obtain the total number of local vectors, and each local vector corresponds to a unique combined index. Among them, the output channel index OC, the input channel group index G, the grid width index u, and the grid height index v together constitute a four-dimensional index group (u, v, OC, G).
[0101] For example, taking an output channel number OC=64, an input channel group number G=4, and a transformation grid width × height of 4×4 with a total of 16 positions as an example, a total of OC×G×16=4096 local weight vectors are generated, and each vector corresponds to a four-dimensional index group.
[0102] In some embodiments, when dividing the local weight vector, one or more dimensions from a variety of dimensions can be selected according to implementation requirements. This application does not impose any limitations on this.
[0103] For each local weight vector, the maximum absolute value is calculated, and this maximum absolute value is set as the weight scale for that local weight vector. Thus, the Winograd domain weight vectors on different output channels, different input channel groups, different Winograd transform grid positions, and different step length branches all have their own independent scale factors.
[0104] It's important to note that, regarding the output channel dimension, different output channels correspond to convolutional kernels that perform different feature extraction functions, with independent dynamic ranges. Setting a scale for each output channel is fundamental to ensuring quantization accuracy. For the input channel group dimension, input channels are divided into groups, sharing a scale within the same group. Regarding the Winograd transform grid position dimension, the coefficients introduced by the transform matrix G differ across rows and columns, resulting in independent dynamic ranges for weight values at different grid positions. Estimating the local maximum absolute value separately for each grid position most accurately reflects the local distribution. Regarding the stride branch dimension, with a stride of 2, the four branches correspond to different subsets of input positions, and the weight distribution of each branch is independent. Setting a scale separately for each branch is crucial to covering its dynamic range.
[0105] The granularity of the structured scale design described above is much finer than that of traditional schemes. It can independently adapt to the dynamic range in each dimension combination of the quantization tensor, ensuring quantization resolution while avoiding the problem of large overall truncation error caused by excessively coarse granularity in traditional schemes.
[0106] After calculating the weight scale, the system quantizes the weight tensor according to the weight scale. The system can perform three quantization operations—scaling, rounding, and clipping—on each local weight vector according to the corresponding weight scale.
[0107] Specifically, it includes the following steps: First, the local weight vector is scaled based on the weight scale corresponding to the local weight vector. Each element in the local weight vector can be scaled by division according to its corresponding weight scale to reduce its numerical range.
[0108] Then, based on the preset bit width, the scaled local weight vector is sequentially rounded and clipped to obtain the quantized weights corresponding to the local weight vector.
[0109] Rounding can be done using either the nearest rounding strategy or the rounding-to-the-top strategy to ensure the accuracy of the subsequent Winograd algorithm.
[0110] The weights are truncated to the effective range boundary by using a preset bit width b_w (e.g., an 8-bit integer int8 or a 4-bit integer int4).
[0111] After completing the above three steps, a quantized weight tensor with a preset bit width format is obtained. Its shape is the same as the Winograd field weight tensor, but the data type changes from FP32 to INT8, reducing storage requirements to 1 / 4 of the original. All elements are within the range that can be directly processed by integer operations. The quantized weight tensor will be stored together with the weight scale tensor for use by the Winograd integer convolution kernel and for inverse scaling during the inference phase.
[0112] Based on the above technical solution, compared with the traditional layer-by-layer quantization or channel-by-channel quantization method, the above vector-by-vector scale quantization method is better adapted to the non-uniform distribution characteristics of Winograd domain weights, more accurately reflects the dynamic range characteristics of each local position of Winograd domain weights, and effectively reduces quantization error.
[0113] In some embodiments, the process of calculating the weight scale of the weight tensor using specific parameter analysis includes: First, verify the dimension and type of the input weight tensor, and then correctly parse the output channel number oc and input channel number ic based on the transpose flag.
[0114] The quantization groups are then divided into num_groups, which consist of 16 input channels, and a floating-point scaling factor tensor of shape [oc, number of input channel groups num_groups] (step 1) or [oc, num_groups, 4] (step 2) is assigned.
[0115] Next, the WinogradWgtFindScaleKernel kernel function is launched. This kernel function executes in parallel with oc×num_groups thread blocks: each thread block is responsible for a combination of an output channel and an input channel group, iterates through the absolute values of all weight elements in the group, and records the maximum value; then, the maximum value is used as the scaling factor and written to the corresponding position of the output tensor; for the case of stride 2, since 4 sub-blocks are generated after the weight transformation (corresponding to the 4 offset regions in the 4×4 output), the kernel function will independently calculate the maximum absolute value for each sub-block, thereby generating a scaling factor of [oc, num_groups, 4].
[0116] In some embodiments, quantization of the weight tensor includes the following steps: The WinogradWgtQuantKernel kernel function is launched, which performs quantization on each weight element: first, it locates the correct scaling factor value based on the step size and transpose flag; then, it calculates the ratio of each weight element to the scale, rounds the ratio, and clips the result to the specified value. Within the range, it is ultimately stored as an int32 type.
[0117] In step S500, the activation scale corresponding to the activation data is calculated, and the activation data is quantized according to the activation scale to obtain quantized activation.
[0118] During the inference process of the JPEGAI network, the distribution of activation data varies significantly across different input channels. Taking a convolutional layer of the synthetic transform network as an example, in its input activation (shape [1, 192, 64, 64], i.e., batch 1, 192 channels, 64×64 resolution), the activation values of some channels (such as channels 0-15) are concentrated in [-0.5, 0.5], while the dynamic range of activation values of other channels (such as channels 64-79) can reach [-3.0, 3.0]. If a uniform scale is adopted across the entire layer, it will inevitably lead to insufficient quantization accuracy in the former or significant truncation in the latter.
[0119] Therefore, this method groups the input activation tensor along the channel dimension. Specifically, it includes the following steps: Obtain the input channel information corresponding to the activated tensor data, including the total number of input channels.
[0120] Based on the input channel information, the activation tensor data is divided into multiple input channel groups, resulting in multiple grouped activation tensors. For example, each group consists of 16 channels. The input activation tensors are divided into groups of 16 channels along the input channel dimension, resulting in a total of C_in / 16 channel groups, where C_in represents the total number of input channels. It should be noted that the values can also be set to 8, 32, or other values depending on the hardware vector width, network structure, and error model; this embodiment does not impose such limitations.
[0121] Obtain the maximum absolute value of the activation tensor data corresponding to multiple input channels within each group activation tensor, and set the maximum absolute value as the activation scale corresponding to each group activation tensor. Each group activation tensor can independently calculate the activation scale, thereby accurately adapting to the dynamic range of channels within each group.
[0122] Specifically, for each group, the maximum absolute value of all elements within the activation tensor is calculated; this maximum absolute value is the activation scale for that group. Each group's activation scale participates independently in subsequent quantization calculations, ensuring dynamic range matching across groups under the same bit-width constraint.
[0123] In some embodiments, after calculating the activation scale, the activation data is quantized according to the activation scale. The quantization operation is similar to the weighted quantization process, and three quantization operations of scaling, rounding, and clipping can be performed sequentially on each grouped activation tensor according to the corresponding activation scale.
[0124] Specifically, it includes the following steps: Based on the activation scale corresponding to each group activation tensor, the activation tensor data corresponding to each group activation tensor is scaled.
[0125] Based on the preset bit width, the scaled activation tensor data is sequentially rounded and pruned to obtain the quantized activation corresponding to the grouped activation tensor.
[0126] Rounding can be done using either the nearest integer or a full rounding strategy to ensure the accuracy of the subsequent Winograd algorithm. The activated values are truncated to the valid range boundary by a preset bit width b_w (e.g., an 8-bit integer int8 or a 4-bit integer int4).
[0127] After performing the above operations on all channel groups one by one, the quantization results of each group are spliced together to restore the original channel order, resulting in the global quantization activation tensor.
[0128] Based on the above technical solution, by scaling and quantizing activations by channel group, the problem of partial channel saturation or insufficient quantization accuracy caused by large differences in the numerical distribution of different input channels is effectively avoided. At the same time, the data layout of the quantized activation tensor is consistent with the input requirements of the Winograd integer convolution kernel, allowing it to directly participate in integer convolution operations within the Winograd domain.
[0129] In some embodiments, for layers with significantly different first-layer convolutions or input dynamic ranges, the system also allows for alternative processing using high-precision integer mapping or approximate pass-through estimators to suppress the backpropagation of first-layer quantization errors.
[0130] In step S600, the weight scale, quantization weight, activation scale, and quantization activation in the Winograd domain are input into the integer convolution kernel in the Winograd domain, and integer convolution operation is performed on the integer convolution kernel to obtain the convolution operation result.
[0131] First, the input quantized tensor can be padded and divided into blocks according to the convolution parameters. Specifically, the quantized activations and quantized weights are divided into blocks according to the Winograd convolution size (e.g., tilesize is 4×4, effective output size is 2×2), and the activations are zero-padding is performed according to the padding parameters to obtain several 4×4 input convolutions.
[0132] In some embodiments, when performing convolution operations on integer convolution kernels, the Winograd algorithm can be used to perform integer convolution operations on the quantized weights and quantized activations in the integer convolution kernels. In this application embodiment, no specific Winograd algorithm is limited; existing Winograd algorithms can be used.
[0133] In step S700, the convolution operation result is reverse-scaled based on the weight scale and the activation scale to obtain restored data in integer format.
[0134] One method for reverse scaling is to multiply the result of the integer convolution operation by the product of the weight scale and the activation scale to obtain the restored data in integer format.
[0135] In step S800, the restored data is superimposed with bias terms to obtain compressed image features in integer format.
[0136] Based on the aforementioned reverse scaling and bias stacking mechanism, this method only uses floating-point multiplication in the final output recovery stage throughout the entire inference chain. The main multiplication and addition operations are entirely handled by efficient integer operations, thus balancing numerical accuracy and computational efficiency.
[0137] Based on the above technical solution, the embodiments of this application have the following technical effects: They realize the transformation from ordinary quantization in the spatial domain to structured quantization in the Winograd domain. This method can non-intrusively replace convolutional layers and transposed convolutional layers without changing the overall network topology, and combine this with a custom computation kernel to complete integer inference. This solution has a clear implementation path, good compatibility, and is user-friendly for engineering developers, significantly reducing inference complexity and improving low-bit deployment performance.
[0138] Corresponding to the aforementioned embodiments of the Winograd-based image compression and quantization method, this application also provides an embodiment of a Winograd-based image compression and quantization apparatus. Figure 2The following are schematic diagrams of the structure of an image compression and quantization device based on Winograd in some embodiments, such as... Figure 2 As shown, the Winograd-based image compression and quantization device 200 includes a parameter acquisition module 210, a convolution kernel parameter acquisition module 220, a weight transformation module 230, a weight quantization module 240, an activation quantization module 250, an integer convolution operation module 260, a restoration module 270, and a bias term superposition module 280.
[0139] The parameter acquisition module is configured to acquire the input parameters of the initial convolution of the target neural network model; the input parameters include activation data and weight data in the original spatial domain; the initial convolution is used to compress and encode the image processed by the target neural network model; The kernel parameter acquisition module is configured to acquire the kernel parameters of the initial convolution; the kernel parameters include the kernel stride. The weight transformation module is configured to transform the weight data from the original spatial domain to the Winograd domain based on the convolution kernel stride, to obtain a weight tensor in the Winograd domain; wherein, when the convolution kernel stride is 1, a first Winograd weight transformation is performed on the weight data to obtain a weight tensor in a first format; when the convolution kernel stride is 2, a second Winograd weight transformation is performed on the weight data to obtain a weight tensor in a second format. The weight quantization module is configured to calculate the weight scale of the weight tensor in the Winograd domain, and to quantize the weight tensor according to the weight scale to obtain the quantized weights in the Winograd domain. The activation quantization module is configured to calculate the activation scale corresponding to the activation data, and to quantize the activation data according to the activation scale to obtain quantized activation. The integer convolution operation module is configured to input the weight scale, the quantization weight, the activation scale, and the quantization activation in the Winograd domain into the integer convolution kernel in the Winograd domain, and perform integer convolution operation on the integer convolution kernel to obtain the convolution operation result. The restoration module is configured to reverse scale the convolution operation result based on the weight scale and the activation scale to obtain restored data in integer format; The bias term overlay module is configured to overlay bias terms on the restored data to obtain compressed image features in integer format.
[0140] Corresponding to the aforementioned embodiments of the Winograd-based image compression and quantization method, this application also provides a computing device, including a processor and a memory. The processor is used to execute a program, specifically performing the relevant steps in the aforementioned embodiments of the Winograd-based image compression and quantization method. The memory is used to store the aforementioned program.
[0141] Some embodiments of this application provide a computer-readable storage medium storing at least one executable instruction that, when executed on a computing device, causes the computing device to perform the Winograd-based image compression and quantization method described in the above embodiments.
[0142] Some embodiments of this application provide a chip system applied to a server. The chip system includes one or more interface circuits and one or more processors. The interface circuits and processors are interconnected via lines. The interface circuits are used to receive signals from the server's memory and send signals to the processors, the signals including computer instructions stored in the memory. When the processor executes the computer instructions, the server performs various steps in the Winograd-based image compression and quantization method shown in the above-described method embodiments.
[0143] The beneficial effects that the readable storage medium provided in some embodiments of this application can achieve can be referred to the beneficial effects of the corresponding Winograd-based image compression and quantization method provided above, and will not be repeated here.
[0144] Similar parts between the embodiments provided in this application can be referred to mutually. The specific implementation methods provided above are only a few examples under the overall concept of this application and do not constitute a limitation on the scope of protection of this application. For those skilled in the art, any other implementation methods extended from the solution of this application without creative effort shall fall within the scope of protection of this application.
Claims
1. An image compression and quantization method based on Winograd, characterized in that, include: Obtain the input parameters for the initial convolution of the target neural network model; The input parameters include activation data and weight data in the original spatial domain; the initial convolution is used to compress and encode the image processed by the target neural network model; Obtain the kernel parameters of the initial convolution; the kernel parameters include the kernel stride; Based on the convolution kernel stride, the weight data is transformed from the original spatial domain to the Winograd domain to obtain a weight tensor in the Winograd domain; wherein, when the convolution kernel stride is 1, a first Winograd weight transformation is performed on the weight data to obtain a weight tensor in a first format; when the convolution kernel stride is 2, a second Winograd weight transformation is performed on the weight data to obtain a weight tensor in a second format. Calculate the weight scale of the weight tensor in the Winograd domain, and quantize the weight tensor according to the weight scale to obtain the quantized weights in the Winograd domain. Calculate the activation scale corresponding to the activation data, and quantize the activation data according to the activation scale to obtain quantized activation; The weight scale, quantization weight, activation scale, and quantization activation in the Winograd domain are input into an integer convolution kernel in the Winograd domain, and an integer convolution operation is performed on the integer convolution kernel to obtain the convolution operation result. Based on the weight scale and the activation scale, the convolution operation result is reverse-scaled to obtain restored data in integer format; The restored data is then superimposed with bias terms to obtain compressed image features in integer format.
2. The image compression and quantization method based on Winograd according to claim 1, characterized in that, The process of obtaining the input parameters for the initial convolution of the target neural network model includes: Determine the initial convolutions to be quantized in the target neural network model; the initial convolutions include ordinary convolutions and transposed convolutions, wherein the ordinary convolutions are convolutions that do not require transposition; Obtain the activation tensor data and convolution kernel weight data corresponding to the initial convolution in the original spatial domain; the activation tensor data and the convolution kernel weight data are in floating-point format.
3. The image compression and quantization method based on Winograd according to claim 2, characterized in that, When the convolution kernel stride is 1, performing a first Winograd weight transformation on the weight data includes: With the convolution kernel stride being 1, each column of the weight data is processed based on the left multiplication transformation function to map each column of data into a four-dimensional intermediate vector; The weighted data is processed based on the right multiplication transformation function to map each row of data into a four-dimensional output vector. A four-row, four-column matrix is generated based on the four-dimensional intermediate vector and the four-dimensional output vector; the four-row, four-column matrix contains a four-dimensional weight tensor; the four dimensions include the output channel dimension, the input channel group dimension, the Winograd transform grid width dimension, and the Winograd transform grid height dimension.
4. The image compression and quantization method based on Winograd according to claim 2, characterized in that, When the convolution kernel stride is 2, performing a second Winograd weight transformation on the weight data includes: When the convolution kernel stride is 2, the convolution kernel corresponding to the weight data is detected; If the convolution kernel is a transposed convolution and the kernel size is 4×4, then the 4×4 weight data is divided into four 2×2 sub-weight data; if the convolution kernel is not a transposed convolution, or the kernel size is not 4×4, then the weight data is divided into four sub-weight data of sizes: 2×2, 2×1, 1×2, and 1×1. The sub-weight data is processed using a left-multiplication transformation function to map each column of data into a five-dimensional intermediate vector; and the sub-weight data is processed using a right-multiplication transformation function to map each row of data into a five-dimensional output vector. A five-dimensional weight tensor is generated based on the five-dimensional intermediate vector and the five-dimensional output vector; the second format includes the output channel dimension, the input channel group dimension, the Winograd transform grid width dimension, the Winograd transform grid height dimension, and the branch dimension.
5. The image compression and quantization method based on Winograd according to claim 3 or 4, characterized in that, The calculation of the weight scale of the weight tensor in the Winograd domain includes: Determine the multiple dimensional information corresponding to the weight tensor under the Winograd domain; Based on the aforementioned multi-dimensional information, the weight tensor is divided into multiple local weight vectors; Calculate the maximum absolute value of each of the multiple local weight vectors, and set the maximum absolute value as the weight scale of the corresponding local weight vector.
6. The image compression and quantization method based on Winograd according to claim 5, characterized in that, The quantization process of the weight tensor according to the weight scale includes: Based on the weight scale corresponding to the local weight vector, the local weight vector is scaled. Based on a preset bit width, the scaled local weight vector is sequentially rounded and clipped to obtain the quantized weights corresponding to the local weight vector.
7. The image compression and quantization method based on Winograd according to claim 2, characterized in that, The calculation of the activation scale corresponding to the activation data includes: Obtain the input channel information corresponding to the activation tensor data; Based on the input channel information, the activation tensor data is divided into multiple input channel groups to obtain multiple grouped activation tensors; Obtain the maximum absolute value of the activation tensor data corresponding to multiple input channels within each group activation tensor, and set the maximum absolute value as the activation scale corresponding to each group activation tensor.
8. The image compression and quantization method based on Winograd according to claim 7, characterized in that, The quantification of the activation data according to the activation scale includes: Based on the activation scale corresponding to each group activation tensor, the activation tensor data corresponding to each group activation tensor is scaled. Based on a preset bit width, the scaled activation tensor data is sequentially rounded and pruned to obtain the quantized activation corresponding to the grouped activation tensor.
9. The image compression and quantization method based on Winograd according to claim 2, characterized in that, The integer convolution operation on the integer convolution kernel includes: The quantization weights and quantization activations in the integer convolution kernel are subjected to integer convolution operations using the Winograd algorithm.
10. An image compression and quantization device based on Winograd, characterized in that, The device includes: The parameter acquisition module is configured to acquire the input parameters of the initial convolution of the target neural network model; the input parameters include activation data and weight data in the original spatial domain; the initial convolution is used to compress and encode the image processed by the target neural network model; The kernel parameter acquisition module is configured to acquire the kernel parameters of the initial convolution; the kernel parameters include the kernel stride. The weight transformation module is configured to transform the weight data from the original spatial domain to the Winograd domain based on the convolution kernel stride, to obtain a weight tensor in the Winograd domain; wherein, when the convolution kernel stride is 1, a first Winograd weight transformation is performed on the weight data to obtain a weight tensor in a first format; when the convolution kernel stride is 2, a second Winograd weight transformation is performed on the weight data to obtain a weight tensor in a second format. The weight quantization module is configured to calculate the weight scale of the weight tensor in the Winograd domain, and to quantize the weight tensor according to the weight scale to obtain the quantized weights in the Winograd domain. The activation quantization module is configured to calculate the activation scale corresponding to the activation data, and to quantize the activation data according to the activation scale to obtain quantized activation. The integer convolution operation module is configured to input the weight scale, the quantization weight, the activation scale, and the quantization activation in the Winograd domain into the integer convolution kernel in the Winograd domain, and perform integer convolution operation on the integer convolution kernel to obtain the convolution operation result. The restoration module is configured to reverse scale the convolution operation result based on the weight scale and the activation scale to obtain restored data in integer format; The bias term overlay module is configured to overlay bias terms on the restored data to obtain compressed image features in integer format.
Citation Information
Patent Citations
Convolution operator capable of accommodating random bit flipping error and application method thereof
CN120975141A
Medical image segmentation method and system based on accelerated convolutional neural network
CN121746402A