Image compression method and system simultaneously facing human eyes and machine vision based on Mama
By introducing Mamba alignment network and knowledge distillation technology in image compression, the problem that the existing technology is difficult to take into account the needs of human eyes and machine vision is solved, and efficient image compression and machine vision tasks are achieved, saving computing resources and improving efficiency.
Patent Information
- Application Number
- CN202510213376.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
AI Technical Summary
Existing image compression methods are difficult to take into account the dual needs of human eye perception and machine vision. Especially in applications such as autonomous driving, intelligent surveillance and medical imaging, traditional methods cannot meet the accuracy needs of machine vision. In edge devices and large-scale real-time machine analysis applications, reconstructing high-quality images requires a large amount of computing resources, resulting in efficiency bottlenecks.
Using Mamba-based image compression method, the potential features are generated through preset analysis and transformation networks, and the hyper-priori entropy model is used for encoding and decoding to generate binary code streams. Then, the potential features are upsampled through the Mamba alignment network, the alignment features are obtained, and input them into the machine vision backend network for knowledge distillation to complete the machine vision task.
It realizes the completion of machine vision tasks without reconstructing images, saves inference time, improves the performance and work efficiency of machine vision tasks, and ensures the high quality of image reconstruction tasks, meeting the dual needs of human eyes and machine vision.
Smart Images

Figure CN120151540A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical fields of image compression and computer vision, and in particular, to an image compression method and system based on Mamba that are simultaneously oriented to human vision and machine vision. Background Art
[0002] With the popularization of intelligent devices and the rapid development of computer vision technology, image data has become an important part of modern communication and computing fields. As an important means to solve the problems of data transmission bandwidth and storage space, image compression technology plays a crucial role. However, most of the existing image compression methods are mainly designed to optimize human visual perception, which leads to many challenges in the application of machine vision and deep learning tasks.
[0003] With the progress of machine vision technology, especially in application fields such as autonomous driving, intelligent monitoring, and medical imaging, the demand for transmitting high-quality visual data across devices is increasing. In these scenarios, visual data not only needs to be viewed by human eyes but also must be analyzed and processed by machine vision systems. Traditional image compression methods first optimize human visual perception, compress and reconstruct images, and then apply them to machine vision tasks, but this method will result in the quality of the reconstructed images not meeting the accuracy requirements of machine vision.
[0004] In addition, in edge devices and large-scale real-time machine analysis applications, reconstructing high-quality images requires a large amount of computing resources, and there are significant bottlenecks in efficiency. Therefore, traditional image compression technologies often have difficulty balancing the dual requirements of human eye perception and machine vision, and it is urgent to explore new solutions to more efficiently address this challenge. Summary of the Invention
[0005] Aiming at the deficiencies in the prior art, the purpose of the present disclosure is to provide an image compression method and system based on Mamba that are simultaneously oriented to human vision and machine vision.
[0006] To achieve the above purpose, according to one aspect of the present disclosure, there is provided an image compression method based on Mamba that is simultaneously oriented to human vision and machine vision, including:
[0007] Input the input RGB image into a preset analysis transformation network to determine potential features;
[0008] Use a preset hyperprior entropy model to sequentially perform entropy parameter estimation, quantization, and entropy coding on the potential features to determine a binary bitstream;
[0009] Use the preset hyperprior entropy model to perform entropy decoding on the binary bitstream to determine the potential features;
[0010] Input the potential feature into a preset synthesis transformation network for inverse transformation to determine a reconstructed image for human eye viewing;
[0011] Perform Mamba upsampling processing on the potential feature to determine an aligned feature;
[0012] Input the aligned feature into a preset machine vision backend network for knowledge distillation to complete a preset machine vision task.
[0013] Optionally, the inputting the input RGB image into a preset analysis transformation network to determine a potential feature includes:
[0014] Perform non-linear transformation processing on the input RGB image to determine an image after non-linear transformation processing;
[0015] Perform downsampling processing on the image after non-linear transformation processing to determine the potential feature.
[0016] Optionally, the preset hyperprior entropy model includes the hyperprior analysis transformation module, the entropy model, and the hyperprior synthesis transformation module.
[0017] Optionally, the binary bitstream includes a binary bitstream containing side information and a binary bitstream containing potential feature information.
[0018] Optionally, the using the preset hyperprior entropy model to sequentially perform entropy parameter estimation, quantization, and entropy coding on the potential feature to determine a binary bitstream includes:
[0019] Input the potential feature into the hyperprior analysis transformation module for entropy parameter estimation to determine side information;
[0020] Perform quantization processing on the side information to determine the quantized side information;
[0021] Input the quantized side information into the entropy model for entropy coding to determine the binary bitstream containing side information;
[0022] Perform quantization processing on the potential feature to determine the quantized potential feature;
[0023] Input the quantized potential feature into the entropy model for entropy coding to determine a binary bitstream containing potential feature information.
[0024] Optionally, the inputting the potential feature into a preset synthesis transformation network for inverse transformation to determine a reconstructed image for human eye viewing includes:
[0025] Perform non-linear transformation processing on the potential feature to determine the potential feature after non-linear transformation processing;
[0026] Upsample the potentially processed features after non-linear transformation to determine the reconstructed image for human eye viewing.
[0027] Optionally, the upsampling process of the potentially processed features based on Mamba to determine the aligned features includes:
[0028] Input the potentially processed features into a preset Mamba alignment network for feature alignment processing to determine the aligned potentially processed features;
[0029] Upsample the aligned potentially processed features to determine the aligned features.
[0030] Optionally, inputting the aligned features into a preset machine vision backend network for knowledge distillation to complete a preset machine vision task includes:
[0031] Determine a teacher network according to the preset machine vision task;
[0032] Train the preset machine vision backend network according to the teacher network to determine the trained machine vision backend network;
[0033] Input the aligned features into the trained machine vision backend network for analysis processing to complete the preset machine vision task.
[0034] According to the second aspect of the present disclosure, there is provided an Mamba-based image compression system for both human eye and machine vision, including:
[0035] An analysis transformation module for inputting an input RGB image into a preset analysis transformation network to determine potentially processed features;
[0036] A hyperprior entropy model for sequentially performing entropy parameter estimation, quantization, and entropy coding on the potentially processed features using a preset hyperprior entropy model to determine a binary bitstream;
[0037] The hyperprior entropy model is further configured to perform entropy decoding on the binary bitstream using the preset hyperprior entropy model to determine the potentially processed features;
[0038] A synthesis transformation module for inputting the potentially processed features into a preset synthesis transformation network for inverse transformation to determine the reconstructed image for human eye viewing;
[0039] An Mamba alignment network module for performing Mamba-based upsampling processing on the potentially processed features to determine aligned features;
[0040] The machine vision task backend module is used to input the aligned features into a preset machine vision backend network for knowledge distillation to complete a preset machine vision task.
[0041] According to a third aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, the steps of the method provided in the first aspect of the present disclosure are implemented.
[0042] According to a fourth aspect of the present disclosure, there is provided an electronic device, including:
[0043] A memory having a computer program stored thereon;
[0044] A processor for executing the computer program in the memory to implement the steps of the method provided in the first aspect of the present disclosure.
[0045] Compared with the prior art, the embodiments of the present disclosure have at least one of the following beneficial effects:
[0046] Through the above technical solution, a compact latent feature with rich semantic information is generated by a preset analysis transformation network, and after encoding and decoding by a preset hyperprior entropy model, on the one hand, the latent feature can compress the input RGB image in a preset synthesis transformation network and reconstruct a reconstructed image for human eyes to view, showing high performance in the image reconstruction task; on the other hand, the latent feature is processed by Mamba upsampling to filter out irrelevant information and obtain aligned features, and the aligned features are input into a preset machine vision backend network to execute a preset machine vision task, without first reconstructing the image and then executing the machine vision task, which can save inference time, improve the performance and work efficiency in the machine vision task, and moreover, knowledge distillation is applied to the preset machine vision backend network to transfer the features in the pixel domain to the compression domain, improving the analysis and processing ability of the machine vision backend network in the compression domain, and by transmitting a single-stream bitstream, high-performance human eye tasks and machine vision tasks are simultaneously realized.
[0047] In the embodiments of the present disclosure, the preset machine vision task can be flexibly extended according to actual needs and will not affect the performance of the image reconstruction task, having flexibility and scalability.
[0048] In the embodiments of the present disclosure, a pixel domain machine vision task network loaded with pre-trained weights with an RGB image as the input is used as the teacher network, and the machine vision task backend network in the compression domain is used as the student network. The student network learns from the teacher network to improve the analysis and processing ability of the student network in the machine vision task in the compression domain and optimize the performance of the student network in the machine vision task. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Other features, objects, and advantages of the present disclosure will become more apparent from the following detailed description of non - limiting embodiments with reference to the accompanying drawings:
[0050] Figure 1 It is a schematic flowchart of an image compression method based on Mamba for both human vision and machine vision according to an exemplary embodiment.
[0051] Figure 2 It is a block diagram of an image compression system based on Mamba for both human vision and machine vision according to an exemplary embodiment. Detailed implementation manners
[0052] The present disclosure will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present disclosure, but do not limit the present disclosure in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present disclosure, several modifications and improvements can still be made. These all fall within the protection scope of the present disclosure.
[0053] The present disclosure provides an image compression method based on Mamba for both human vision and machine vision, which focuses on visual data applied to human vision tasks and machine vision tasks, and can simultaneously complete high - quality image reconstruction tasks and high - performance machine analysis tasks.
[0054] Figure 1 It is a schematic flowchart of an image compression method based on Mamba for both human vision and machine vision according to an exemplary embodiment.
[0055] As Figure 1 shown, an image compression method based on Mamba for both human vision and machine vision includes S11 to S16.
[0056] S11, input the input RGB image into a preset analysis transformation network to determine latent features.
[0057] The preset analysis transformation network includes a non - linear transformation and a 16 - fold downsampling.
[0058] Among them, the latent features have rich semantic information and are compact.
[0059] S12, use a preset hyper - prior entropy model to sequentially perform entropy parameter estimation, quantization, and entropy coding on the latent features to determine a binary bitstream.
[0060] Among them, the binary bitstream is used for storage and transmission.
[0061] As Figure 1 shown, steps S11 and S12 are executed at the encoding end.
[0062] S13. Entropy decode the binary bitstream using a preset hyperprior entropy model to determine latent features.
[0063] S14. Input the latent features into a preset synthesis transformation network for inverse transformation to determine a reconstructed image for human eye viewing.
[0064] S15. Perform Mamba upsampling processing on the latent features to determine aligned features.
[0065] Among them, the aligned features contain rich semantics.
[0066] S16. Input the aligned features into a preset machine vision backend network for knowledge distillation to complete a preset machine vision task.
[0067] Such as Figure 1 As shown, the above steps S13 to S16 are executed at the decoding end.
[0068] The method provided by the present disclosure does not require fine-tuning the encoder. At the decoding end, compact latent features are decoded from the binary bitstream. By transmitting a single bitstream, high-performance human eye tasks and machine vision tasks can be achieved simultaneously, reducing training costs and storage overhead.
[0069] Through the above technical solution, a preset analysis transformation network is used to generate compact latent features with rich semantic information. After encoding and decoding through a preset hyperprior entropy model, on the one hand, the latent features can compress the input RGB image in a preset synthesis transformation network and reconstruct a reconstructed image for human eye viewing, showing high performance in the image reconstruction task; on the other hand, perform Mamba upsampling processing on the latent features to filter out irrelevant information and obtain aligned features. Input the aligned features into a preset machine vision backend network to execute a preset machine vision task, without first reconstructing the image and then executing the machine vision task, which can save inference time, improve the performance and working efficiency in the machine vision task. Moreover, apply knowledge distillation on the preset machine vision backend network to transfer the features in the pixel domain to the compression domain, improving the analysis and processing ability of the machine vision backend network in the compression domain. By transmitting a single bitstream, high-performance human eye tasks and machine vision tasks can be achieved simultaneously.
[0070] In a possible embodiment, S11. Input the input RGB image into a preset analysis transformation network to determine latent features, which may include S111 to S112.
[0071] S111. Perform non-linear transformation processing on the input RGB image to determine the image after non-linear transformation processing.
[0072] S112. Downsample the non-linearly transformed image to determine the latent features.
[0073] Among them, the downsampling process can adopt 16-fold downsampling.
[0074] As an example, first perform non-linear transformation on the RGB image x with size h×w to determine the non-linearly transformed image, and then perform 16-fold downsampling on the non-linearly transformed image to convert it to a size of:
[0075] The latent feature y with rich semantic information and compactness.
[0076] In a possible embodiment, the preset hyperprior entropy model includes a hyperprior analysis transformation module h a , an entropy model, and a hyperprior synthesis transformation module h s .
[0077] Among them, the entropy model includes two processes: entropy encoding AE and entropy decoding AD.
[0078] In a possible embodiment, S12. Use the preset hyperprior entropy model to sequentially perform entropy parameter estimation, quantization, and entropy encoding on the latent features to determine the binary bitstream, including S121 to S125.
[0079] Among them, the binary bitstream includes a binary bitstream containing side information and a binary bitstream containing latent feature information.
[0080] S121. Input the latent features into the hyperprior analysis transformation module for entropy parameter estimation to determine the side information.
[0081] Among them, the process of entropy parameter estimation is to model the latent features as a Gaussian distribution to obtain the side information of the probability model containing the latent features, which is convenient for compression encoding of the latent features.
[0082] S122. Quantize the side information to determine the quantized side information.
[0083] S123. Input the quantized side information into the entropy model for entropy encoding to determine the binary bitstream containing the side information.
[0084] As an example, the latent feature y is input into the hyperprior analysis transformation module to obtain the side information z, and the side information z is quantized Q to obtain the quantized side information Use the entropy model to Perform entropy encoding on the quantized side information to form a binary bitstream, that is, the binary bitstream containing the side information.
[0085] S124. Quantize the potential features to determine the quantized potential features.
[0086] S125. Input the quantized potential features into an entropy model for entropy encoding to determine a binary bitstream containing potential feature information.
[0087] As another example, quantize the potential feature y as Q to obtain the quantized potential feature Use the hyperprior entropy model for the quantized potential feature to perform entropy encoding to form a binary bitstream, that is, the binary bitstream containing the information of the potential feature.
[0088] Store or transmit both of the above two binary bitstreams, that is, the binary bitstream containing side information and the binary bitstream containing potential feature information.
[0089] In a possible embodiment S13, use a preset hyperprior entropy model to perform entropy decoding on the binary bitstream to determine the potential features, including:
[0090] Use the entropy model in the hyperprior entropy model to perform entropy decoding on the binary bitstream containing side information to obtain the quantized side information Input the quantized side information into the hyperprior synthesis transformation module h s to obtain μ and σ containing potential feature probability information 2 ; according to μ and σ containing potential feature probability information 2 , use the entropy model to perform entropy decoding on the binary bitstream containing the potential feature to obtain the compact potential feature y.
[0091] In the present disclosure, the above steps S12 to S13 can be expressed by the following formula:
[0092] z = h a (y)
[0093]
[0094] where z represents side information, h a represents the hyperprior analysis transformation module, y represents the potential feature, represents the quantized side information, AD represents entropy decoding, AE represents entropy encoding, Q represents quantization, μ, σ 2 represent the potential feature probability information, h s represents the hyperprior synthesis transformation module, represents the quantized potential feature, N represents the Gaussian probability distribution, represents the probability of the quantized potential feature of.
[0095] In a possible embodiment, in S14, the latent features are input into a preset synthesis transformation network for inverse transformation to determine a reconstructed image for human eye viewing, including: S141 to S142.
[0096] Among them, the preset synthesis transformation network includes a non-linear transformation and a 16-fold upsampling process to obtain the reconstructed image for the human eye viewing task.
[0097] S141, perform a non-linear transformation process on the latent features to determine the latent features after the non-linear transformation process.
[0098] S142, perform an upsampling process on the latent features after the non-linear transformation process to determine the reconstructed image for human eye viewing.
[0099] As an example, the latent features with a size of:
[0100] are transformed into an RGB image with a size of h×w through a non-linear transformation and a 16-fold upsampling process for the human eye viewing task.
[0101] In the present disclosure, the network models involved in steps S11 to S14 are all designed for human perception, such as the preset analysis transformation network, the preset hyperprior entropy model, and the preset synthesis transformation network, and the following rate-distortion function can be used for optimization:
[0102] Loss = R + λD
[0103] Among them, R represents the rate estimation loss of the latent features and the side information, D represents the mse loss between the input RGB image and the reconstructed image and λ represents the balance factor.
[0104] The balance factor λ is used as the rate estimation loss R of the latent features and the side information and the mse loss D between the input RGB image and the reconstructed image to adjust the bit rate.
[0105] In a possible embodiment, in S15, perform an upsampling process on the latent features based on Mamba to determine the alignment features, including: S151 to S152.
[0106] S151, input the latent features into a preset Mamba alignment network for feature alignment processing to determine the aligned latent features.
[0107] Among them, the preset Mamba alignment network includes a Mamba block and a PixelShuffle 4-fold upsampling.
[0108] The potential features are derived from steps S11 to S14.
[0109] S152. Upsample the aligned potential features to determine the aligned features.
[0110] Among them, in this step, the upsampling process is 4-fold upsampling, and the aligned features have rich semantic information.
[0111] As an example, the potential features with the size in the compressed domain:
[0112] are subjected to feature alignment processing through the Mamba block to determine the aligned potential features, and the aligned potential features are upsampled 4-fold through PixelShuffle 4-fold upsampling, so as to obtain the aligned features with the size:
[0113]
[0114] The preset Mamba alignment network can use the selection mechanism of Mamba to filter out the features irrelevant to the downstream machine vision tasks and retain the necessary and relevant features, and match the size of the input feature tensor of the machine vision task network, which is convenient for subsequent machine analysis and processing.
[0115] In a possible embodiment, S16. Input the aligned features into a preset machine vision backend network for knowledge distillation to complete the preset machine vision task, which may include: S161 to S163.
[0116] S161. Determine the teacher network according to the preset machine vision task.
[0117] Among them, the teacher network is a machine vision network based on MambaVision.
[0118] S162. Train the model of the preset machine vision backend network according to the teacher network to determine the trained machine vision backend network.
[0119] Among them, the preset machine vision backend network is a machine vision network based on MambaVision removing the Stem processing module in the head of MambaVision, retaining the stage1, stage2, stage3, and stage4 modules and the network of the machine vision task detector as the student network.
[0120] Knowledge distillation represents the process in which the student network trains the model by learning the output of the teacher network and combining the machine vision task loss.
[0121] Use the pre-trained complete MambaVision-based machine vision network as the teacher network, and use the preset machine vision backend network as the student network. The student network is trained by learning the outputs of stages 1 to 4 of the teacher network and combining the machine vision task loss.
[0122] S163. Input the aligned features into the trained machine vision backend network for analysis and processing to complete the preset machine vision task.
[0123] As an example, for image classification as the preset machine vision task, select MambaVision as the image classification analysis network and use it as the teacher network. Remove the processing module at the head of MambaVision, and retain the stage1, stage2, stage3, and stage4 modules and the classifier as the student network, that is, the preset machine vision backend network. Among them, the student network is an image classification network.
[0124] Pass the original RGB image in the pixel domain through stage1, stage2, stage3, and stage4 in sequence to obtain its intermediate feature outputs f' 1 、f' 2 、f' 3 、f' 4 .
[0125] Input the aligned features in the compressed domain into the student network, and pass through stage1, stage2, stage3, and stage4 in sequence to obtain its intermediate feature outputs f 1 、f 2 、f 3 、f 4 .
[0126] In the present disclosure, the networks involved in steps S11 to S14 have been optimized through the rate-distortion function, and the network weights have been frozen in subsequent network optimizations to ensure that the quality of image reconstruction is not affected. For the networks involved in S15 to S16, which are related to the implemented machine vision tasks, knowledge distillation is adopted and optimized through the following loss function:
[0127] Loss = L task + βL teach
[0128] L teach = d(f 1 , f′ 1 ) + d(f 2 , f′ 2 ) + d(f 3 , f′ 3 ) + d(f 4 , f′4 )
[0129] Among them, L task represents the loss of the optimized pixel-domain machine vision task network, and L teach represents the loss of knowledge distillation, that is, the distance loss between the output of the pixel-domain teacher network and the output of the compressed-domain machine vision backend network, and β represents the balance factor.
[0130] β represents the loss L of the pixel-domain machine vision task network task and the loss L of knowledge distillation teach between the balance factor.
[0131] As an example, the loss of the image classification network is the cross-entropy loss, the mse is used to measure the distance loss, and the value of β is 1.
[0132] After the preset machine vision backend network is trained by knowledge distillation, the aligned features in the compressed domain are input into the trained machine vision backend network for analysis and processing, and the image classification task is executed and completed.
[0133] The present disclosure uses the pixel-domain machine vision task network loaded with pre-trained weights taking the RGB image as the input as the teacher network, and the machine vision task backend network in the compressed domain as the student network. The student network learns from the teacher network to improve the analysis and processing capabilities of the student network in the machine vision task in the compressed domain, and optimizes the performance of the student network in the machine vision task.
[0134] In a possible embodiment, the preset machine vision task can also be set to machine analysis tasks such as object detection and semantic segmentation. Its specific vision network can also adopt the resnet50 backbone network, and the machine vision task is implemented based on the above steps S15 to S16. And each preset machine vision task corresponds to an alignment network and a machine vision backend network to implement the preset machine vision task. In the extended machine vision task, only the corresponding alignment network and machine vision backend network need to be optimized.
[0135] The preset machine vision task can be flexibly extended according to actual needs, and does not affect the performance of the image reconstruction task, and has flexibility and scalability.
[0136] The above embodiments of the present disclosure are based on the compact latent features generated by optimizing human perception, which can not only implement the high-quality image reconstruction task for human eye vision through step S14, but also simultaneously implement the high-performance machine vision task through steps S15 to S16, and have flexibility and scalability, that is, transmitting a single-stream latent feature can simultaneously implement multiple vision tasks, and the machine vision task does not need to reconstruct the image and then perform machine analysis, reducing the model complexity and greatly saving the inference time.
[0137] An image compression method based on Mamba for both human eyes and machine vision provided by the present disclosure uses an image classification task as a machine vision task, and evaluates the performance of the method provided by the present disclosure in image reconstruction and image classification through experiments on the public dataset ImageNet.
[0138] The size of the input RGB image is adjusted to 224×224, and this experiment can be implemented using the Pytorch framework on the Win10 system.
[0139] Table 1 shows the performance of image reconstruction using the method provided by the present disclosure at four bitrate points.
[0140] bpp psnr ms-ssim 0.1415 28.45dB 0.9432 0.2111 29.78dB 0.9597 0.3011 31.10dB 0.9715 0.4542 33.03dB 0.9816
[0141] Table 1
[0142] As shown in Table 1, bpp represents the bitrate, psnr represents the Peak Signal-to-Noise Ratio (PSNR), and ms-ssim represents the Multi-Scale Structural Similarity Index (MS-SSIM). Both are performance metrics for evaluating the image reconstruction task. The larger the values of psnr and ms-ssim, the better the image reconstruction quality.
[0143] Table 2 shows the comparison of Top-1 accuracy of the method provided by the present disclosure with Comparative Method 1 and Comparative Method 2 at the same bitrate.
[0144] 0.1415bpp 0.2111bpp 0.3011bpp 0.4542bpp This method 76.45% 77.74% 78.36% 79.52% Comparison method 1 76.07% 77.45% 78.06% 79.26% Comparison method 2 52.61% 60.14% 67.16% 73.84%
[0145] Table 2
[0146] Among them, Comparative Method 1 is a method from the 2023 TCSVT unified architecture adaptation work, which replaces the preset machine vision backend network with the MambaVision image classification network. Comparative Method 2 is a traditional method, that is, after reconstructing the latent features in the compression domain into an RGB image, the reconstructed image is input into the pre-trained MambaVision image classification network for analysis and processing to perform the image classification task. The method provided by the present disclosure, Comparative Method 1, and Comparative Method 2 all use the same learning-based image compression network. Thus, the image reconstruction performance in Table 1 can be achieved at the same bitrate.
[0147] It can be seen from Table 2 that the accuracy of the method provided by the present disclosure in performing the image classification task is better than that of Comparative Method 1 and Comparative Method 2.
[0148] Table 3 shows the comparison of the model complexity between the method provided by the present disclosure and the traditional method at the decoding end (excluding entropy decoding) when performing an image classification task.
[0149] FLOPs Params. Inference time (GPU) This method 4.925G 33.160M 0.2ms Traditional method 32.174G 43.415M 0.3ms
[0150] Table 3
[0151] As shown in Table 3, FLOPs represents the number of floating-point operations during the model's task execution, Params. represents the number of parameters in the model, and the inference time at the decoding end is obtained through experiments on a 4090 GPU. Compared with the traditional method, the method provided by the present disclosure can greatly reduce the model complexity and the inference time.
[0152] Figure 2 It is a block diagram of an image compression system based on Mamba that is simultaneously oriented to human vision and machine vision shown according to an exemplary embodiment.
[0153] Based on the same concept, the present disclosure also provides an image compression system 100 based on Mamba that is simultaneously oriented to human vision and machine vision, as Figure 2 shown, including: an analysis transformation module 110, a hyperprior entropy model 120, a synthesis transformation module 130, a Mamba alignment network module 140, and a Mamba alignment network module 150.
[0154] The analysis transformation module 110 is configured to input the input RGB image into a preset analysis transformation network to determine latent features;
[0155] The hyperprior entropy model 120 is configured to sequentially perform entropy parameter estimation, quantization, and entropy coding on the latent features using a preset hyperprior entropy model to determine a binary bitstream;
[0156] The hyperprior entropy model 120 is also configured to perform entropy decoding on the binary bitstream using a preset hyperprior entropy model to determine latent features;
[0157] The synthesis transformation module 130 is configured to input the latent features into a preset synthesis transformation network for inverse transformation to determine a reconstructed image for human eye viewing;
[0158] The Mamba alignment network module 140 is configured to perform Mamba upsampling processing on the latent features to determine aligned features;
[0159] The machine vision task backend module 150 is configured to input the aligned features into a preset machine vision backend network for knowledge distillation to complete a preset machine vision task.
[0160] Through the above technical solutions, a compact latent feature with rich semantic information is generated by a preset analysis transformation network, and after encoding and decoding through a preset hyperprior entropy model, on the one hand, the latent feature can compress the input RGB image in a preset synthesis transformation network and reconstruct a reconstructed image for human eyes to view, showing high performance in the image reconstruction task; on the other hand, the latent feature is processed by Mamba upsampling to filter out irrelevant information and obtain aligned features, and the aligned features are input into a preset machine vision backend network to perform a preset machine vision task. There is no need to reconstruct the image first and then perform the machine vision task, which can save inference time, improve the performance and working efficiency in the machine vision task. Moreover, knowledge distillation is applied to the preset machine vision backend network to transfer the features in the pixel domain to the compression domain, improving the analysis and processing ability of the machine vision backend network in the compression domain. By transmitting a single stream of bitstream, high-performance human eye tasks and machine vision tasks can be achieved simultaneously.
[0161] Regarding the embodiments of the above system, the specific ways in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0162] Based on the same inventive concept, in another embodiment of the present disclosure, an electronic device is further provided, including a memory, a processor, and a computer program stored on the memory and capable of running on the processor. When the processor executes the program, it is used to execute an image compression method based on Mamba that is simultaneously oriented to human eyes and machine vision.
[0163] Optionally, the memory is used to store programs; the memory may include volatile memory (English: volatile memory), such as random access memory (English: random-access memory, abbreviation: RAM), such as static random access memory (English: static random-access memory, abbreviation: SRAM), double data rate synchronous dynamic random access memory (English: Double Data Rate Synchronous Dynamic Random Access Memory, abbreviation: DDR SDRAM), etc.; the memory may also include non-volatile memory (English: non-volatile memory), such as flash memory (English: flash memory). The memory is used to store computer programs (such as application programs and functional modules for implementing the above method), computer instructions, etc. The above computer programs, computer instructions, etc. can be stored in partitions in one or more memories. And the above computer programs, computer instructions, data, etc. can be called by the processor.
[0164] The above-mentioned computer programs, computer instructions, etc. can be stored in partitions in one or more memories. And the above-mentioned computer programs, computer instructions, data, etc. can be called by the processor.
[0165] A processor, configured to execute the computer program stored in the memory to implement each step in the method related to the above-mentioned embodiments. For specific details, reference can be made to the relevant descriptions in the foregoing method embodiments.
[0166] The processor and the memory can be of independent structures or integrated structures. When the processor and the memory are of independent structures, the memory and the processor can be coupled and connected through a bus.
[0167] In an embodiment of the present disclosure, there is also provided a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of an image compression method based on Mamba and simultaneously for human eyes and machine vision in any of the above embodiments.
[0168] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, a system, or a computer program product. Therefore, the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program code.
[0169] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0170] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0171] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for realizing the functions specified in one process or a plurality of processes and / or boxes Figure 1 one process or a plurality of processes and / or boxes Figure 1 or steps for realizing the functions specified in a plurality of boxes.
[0172] Although the preferred embodiments of the present disclosure have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic creative concept. Therefore, the appended claims are intended to be construed to cover the preferred embodiments as well as all changes and modifications falling within the scope of the present disclosure.
[0173] Obviously, those skilled in the art can make various changes and modifications to the present disclosure without departing from the spirit and scope of the present disclosure. Thus, if these modifications and variations of the present disclosure fall within the scope of the claims of the present disclosure and their equivalent technologies, the present disclosure is also intended to include these modifications and variations.
Claims
1. A Mamba-based image compression method for both human eyes and machine vision, characterized in that: include: The input RGB image is fed into a preset analysis transformation network to determine the latent features; Using a preset hyper-prior entropy model to sequentially perform entropy parameter estimation, quantization and entropy coding on the potential features to determine a binary code stream; Performing entropy decoding on the binary code stream using the preset hyper-prior entropy model to determine the potential features; Inputting the latent features into a preset synthetic transformation network for inverse transformation to determine a reconstructed image for human eyes to view; Performing Mamba-based upsampling processing on the potential features to determine alignment features; The alignment features are input into a preset machine vision backend network for knowledge distillation to complete the preset machine vision task.
2. The method according to claim 1, characterized in that The input RGB image is input into a preset analysis transformation network to determine potential features, including: Performing nonlinear transformation processing on the input RGB image to determine an image processed by the nonlinear transformation; Down-sampling is performed on the image that has been processed by the nonlinear transformation to determine the potential features.
3. The method according to claim 1, characterized in that The preset super a priori entropy model includes the super a priori analysis transformation module, the entropy model and the super a priori synthesis transformation module.
4. The method according to claim 3, characterized in that The binary code stream includes a binary code stream containing side information and a binary code stream containing potential feature information; The method of using a preset hyper-prior entropy model to sequentially perform entropy parameter estimation, quantization and entropy coding on the potential features to determine a binary code stream includes: Inputting the potential features into the super-prior analysis transformation module to perform entropy parameter estimation and determine side information; quantizing the side information to determine quantized side information; Inputting the quantized side information into the entropy model for entropy coding to determine the binary code stream containing the side information; Quantifying the potential features to determine quantified potential features; The quantized potential features are input into the entropy model for entropy coding to determine a binary code stream containing potential feature information.
5. The method according to claim 1, characterized in that The step of inputting the potential features into a preset synthetic transformation network for inverse transformation to determine a reconstructed image for human eyes to view includes: Performing nonlinear transformation processing on the latent features to determine the latent features after the nonlinear transformation processing; The latent features processed by the nonlinear transformation are up-sampled to determine the reconstructed image for human eyes to view.
6. The method according to claim 1, characterized in that The performing Mamba-based upsampling processing on the potential features to determine the alignment features includes: Inputting the potential features into a preset Mamba alignment network for feature alignment processing to determine the aligned potential features; An upsampling process is performed on the aligned latent features to determine the aligned features.
7. The method according to claim 1, characterized in that The step of inputting the alignment features into a preset machine vision backend network for knowledge distillation to complete a preset machine vision task includes: Determine the teacher network according to the preset machine vision task; Performing model training on the preset machine vision backend network according to the teacher network to determine a trained machine vision backend network; The alignment features are input into the trained machine vision backend network for analysis and processing to complete the preset machine vision task.
8. An image compression system based on Mamba for both human eye and machine vision, characterized in that: include: An analysis and transformation module, which is used to input the input RGB image into a preset analysis and transformation network to determine the potential features; A super a priori entropy model is used to sequentially perform entropy parameter estimation, quantization and entropy coding on the potential features using a preset super a priori entropy model to determine a binary code stream; The super a priori entropy model is also used to perform entropy decoding on the binary code stream using the preset super a priori entropy model to determine the potential features; A synthesis transformation module, used for inputting the potential features into a preset synthesis transformation network for inverse transformation to determine a reconstructed image for human viewing; A Mamba alignment network module, used for performing Mamba upsampling processing on the latent features to determine alignment features; The machine vision task backend module is used to input the alignment features into a preset machine vision backend network for knowledge distillation to complete the preset machine vision task.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.
10. An electronic device, characterized in that: include: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of claims 1 to 7.
Citation Information
Cited By
Man-machine cooperation continuous image compression method and system based on hybrid expert adapter
CN122002039A