Image compression method and device, equipment and storage medium
By extracting and compressing the common and unique features of multiple images and fusing them in the feature domain, the problem of information redundancy in image encoding is solved, and more efficient image compression is achieved.
Patent Information
- Application Number
- CN202510111068.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-30
AI Technical Summary
Multiple images captured in the same scenario contain similarity, resulting in information redundancy during encoding, increasing the pressure for storage and transmission.
The common features and unique features of multiple images to be compressed are extracted, and these features are compressed using the conditional cross-modal entropy model, and fused in the feature domain to avoid repeated compression of common features.
By avoiding repeated compression of common features in multiple images, the encoding bit rate is saved, the storage and transmission pressure is reduced, and the image compression efficiency is improved.
Smart Images

Figure CN120071033A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image processing, and in particular, to an image compression method, apparatus, device, and storage medium. Background Art
[0002] In order to store and transmit images conveniently, image compression is very necessary. The goal of image compression is to minimize the coding bit rate while maintaining the image quality.
[0003] For two images captured in different bands of the same scene, similarities are expected to be included in the spatial or feature domain. If these images are encoded separately, it may lead to significant information redundancy, thus increasing the pressure of storage and transmission. Summary of the Invention
[0004] In order to solve the above technical problems, the present application provides an image compression method, apparatus, device, and storage medium, which avoid repeated compression of common features in multiple images, save the coding bit rate, and reduce the pressure of storage and transmission.
[0005] In a first aspect, the present application provides an image compression method, which includes: extracting the common features of multiple images to be compressed and the unique features of each image to be compressed, where the multiple images to be compressed are images obtained in different bands for the same scene; using a conditional cross-modal entropy model to compress the common features and the unique features of each image to be compressed, to obtain the compressed common features and the compressed unique features of each image to be compressed; and fusing the compressed common features with the compressed unique features of each image to be compressed respectively, to obtain the compressed images corresponding to each image to be compressed.
[0006] In a second aspect, the present application provides an image compression apparatus, which includes: a feature extraction module, configured to extract the common features of multiple images to be compressed and the unique features of each image to be compressed, where the multiple images to be compressed are images obtained in different bands for the same scene; a feature compression module, configured to use a conditional cross-modal entropy model to compress the common features and the unique features of each image to be compressed, to obtain the compressed common features and the compressed unique features of each image to be compressed; and a feature fusion module, configured to fuse the compressed common features with the compressed unique features of each image to be compressed respectively, to obtain the compressed images corresponding to each image to be compressed.
[0007] In a third aspect, the present application provides an electronic device, which includes an image compression device, and the electronic device includes: one or more processors; a storage device, configured to store one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the image compression method as described in the first aspect above.
[0008] Fourthly, the present application provides a storage medium, which may be a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the image compression method in the first aspect above is implemented.
[0009] Fifthly, an embodiment of the present application provides a computer program product, which includes a computer program or instruction. When the computer program or instruction is executed by a processor, any one of the image compression methods in the first aspect above is implemented.
[0010] An embodiment of the present application provides an image compression method, device, equipment and storage medium. The method includes: extracting the common features of multiple images to be compressed and the unique features of each image to be compressed, where the multiple images to be compressed are images obtained in different bands for the same scene; using a conditional cross-modal entropy model to compress the common features and the unique features of each image to be compressed, so as to obtain the compressed common features and the compressed unique features of each image to be compressed; fusing the compressed common features with the compressed unique features of each image to be compressed respectively to obtain the compressed images corresponding to each image to be compressed. In the embodiment of the present application, the features in each image to be compressed are divided into two parts, that is, the common features shared by multiple images to be compressed and the unique features exclusive to each image to be compressed. The common and unique features are compressed respectively, and then fused in the feature domain to reconstruct the compressed image. Since the common features are encoded only once, the repeated compression of the common features in multiple images is avoided, the coding bit rate is saved, and the pressure of storage and transmission is reduced. Description of the Drawings
[0011] The drawings here are incorporated into the specification and constitute a part of this specification, showing the embodiments in line with the present application, and are used together with the specification to explain the principles of the present application.
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained according to these drawings without creative efforts.
[0013] Figure 1 It is a schematic flow chart of the image compression method provided by the embodiment of the present application;
[0014] Figure 2 It is a schematic structural diagram of the DeepRNC network architecture provided by the embodiment of the present application;
[0015] Figure 3 It is a schematic flow chart of extracting image features provided by the embodiment of the present application;
[0016] Figure 4 Schematic diagram of the structure of the convolutional sparse coding block provided by the embodiment of the present application;
[0017] Figure 5 Schematic diagram of the structure of the CFD module structure provided by the embodiment of the present application;
[0018] Figure 6 Schematic diagram of the process of image feature compression provided by the embodiment of the present application;
[0019] Figure 7 Schematic diagram of the structure of the CCE module provided by the embodiment of the present application;
[0020] Figure 8 Schematic diagram of the process of image feature fusion provided by the embodiment of the present application;
[0021] Figure 9 Schematic diagram of the structure of the feature fusion module provided by the embodiment of the present application;
[0022] Figure 10 Schematic diagram of the structure of the image compression device provided by the embodiment of the present application;
[0023] Figure 11 Schematic diagram of the structure of the electronic device provided by the embodiment of the present application. Detailed implementation manners
[0024] In order to be able to more clearly understand the above objects, features and advantages of the present application, the solutions of the present application will be further described below. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other.
[0025] Many specific details are set forth in the following description in order to provide a thorough understanding of the present application, but the present application may be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present application, rather than all of the embodiments.
[0026] The term "including" and its variations used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0027] It should be noted that the concepts such as "first" and "second" mentioned in the present application are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or the interdependence relationship.
[0028] It should be noted that the modifiers "one" and "multiple" mentioned in this application are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly specified in the context, it should be understood as "one or more".
[0029] The following will combine the accompanying drawings and specific implementation manners to elaborate in detail on the image compression method provided by the implementation of this application.
[0030] Figure 1 It is a flowchart of an image compression method in an embodiment of this application. This embodiment is applicable to the situation of compressing multiple images simultaneously. This method can be executed by an image compression device, which can be implemented in software and / or hardware, and the image compression device can be configured in an electronic device.
[0031] As Figure 1 shown, the image compression method provided by the embodiment of this application mainly includes steps S101 - S103.
[0032] S101. Extract the common features of multiple images to be compressed and the unique features of each image to be compressed, where the multiple images to be compressed are images obtained in different bands for the same scene.
[0033] An image to be compressed refers to a digital image file that is planned or needs to be compressed. Compressing an image refers to the process of reducing the size of an image file through a specific algorithm, so as to save storage space, speed up the transmission speed, or adapt to certain applications with limitations on file size.
[0034] Image compression mainly includes two major types of compression methods: lossy compression and lossless compression. Lossy compression removes some image data to reduce the size of the image file, which may lead to a decrease in image quality. Lossless compression does not lose any image information, so the image quality remains unchanged after decompression. The image compression method in the embodiment of this application is a lossy compression method.
[0035] A band refers to a specific wavelength range of electromagnetic radiation captured by a sensor. Each band corresponds to a narrow wavelength range. Further, the bands include: visible light band, near-infrared band, short-wave infrared band, thermal infrared band, microwave / radar band, etc. Images obtained in different bands refer to images captured by a sensor in different parts of the electromagnetic spectrum (i.e., different wavelength ranges or bands). Each band corresponds to a specific spectral region, and these images can reveal different information about the observed object.
[0036] The multiple images to be compressed are images obtained in different bands for the same scene, which means that the multiple images to be compressed are captured by sensors in the same geographical area or scene in different bands of the electromagnetic spectrum (i.e., different spectral ranges). The images in each band reflect the radiation characteristics of the scene within a specific wavelength range and can thus provide different information about the scene.
[0037] Common features refer to the characteristics or attributes shared by all the images to be compressed. Common features ensure the compatibility and consistency among the images. Exclusive features, on the other hand, refer to the attributes or information unique to each image to be compressed, which reflect the uniqueness of each image to be compressed.
[0038] Furthermore, the multiple images to be compressed are captured by different sensors in the same scene. Therefore, the multiple images to be compressed from different sensors will share some common features, while each image to be compressed will retain its exclusive features. Thus, a method of decoupling common and exclusive features can be used for the recovery and fusion of multiple images to be compressed, which can improve the accuracy of recovery and fusion. Further, the multiple images to be compressed are captured by different sensors in the same scene. The multiple images to be compressed can share common semantic and edge information while differing in brightness and texture details. Therefore, by avoiding repeated compression of common features, the common-exclusive feature segmentation strategy can effectively improve the compression efficiency of multiple images to be compressed.
[0039] In a possible implementation, the embodiments of this application utilize a Deep Recurrent Neural Coding (DeepRNC) network architecture, as Figure 2 shown. The DeepRNC network architecture provided in the embodiments of this application mainly includes a Common Feature Disentanglement (CFD) module 21, a Triplebranch Cross-modal Compression (TCC) module 22, and a Multi-modal Feature Fusion (MFF) module 23. It should be noted that in the embodiments of this application, it is exemplified by multiple images to be compressed including a visible light image x 1 and a near-infrared image x 2 for illustration.
[0040] As Figure 2 shown, the input data in the DeepRNC network architecture is the visible light image x 1 and the near-infrared image x 2, the Common Feature Disentanglement (CFD) module 21 is the input module of the DeepRNC network architecture, that is, the input data of the CFD module 21 is the visible light image x 1 and the near-infrared image x 2 . The role of the CFD module 21 is to separate the common feature g 1 between the visible light image x 2 and the near-infrared image x C . The expression of the common feature g C is shown in formula (1).
[0041] g C = f CFD (x 1 , x 2 )(1)
[0042] where f CFD represents the function of the CFD module.
[0043] Then, the common feature g C is processed through two convolutional layers to obtain the common part image f 1 of the visible light image x 2 and the near-infrared image x C1 (g C ). The unique feature of the visible light image x 1 can be obtained by subtracting the common feature g 1 from the visible light image x C . The image x 1 of the unique part of the visible light image x 1u can be represented by formula (2).
[0044] x 1u = x 1 - f C1 (g C )(2)
[0045] where, f C1 represents the function of the convolutional layer, f C1 (g C ) is the common part image of the visible light image x 1 and the near-infrared image x 2 . x 1u is the image of the unique part of the visible light image x 1 .
[0046] The unique feature of the near-infrared image x 2 can be obtained by subtracting the common feature g 2 from the near-infrared image x C to obtain the near-infrared image x2 The image x of the unique part 2u can be represented by formula (3).
[0047] x 2u = x 2 - f C2 (g C )(3)
[0048] where f C2 represents the function of the convolutional layer, and f C2 (g C ) is the common part image of the visible light image x 1 and the near-infrared image x 2 . x 2u is the image of the unique part of the near-infrared image x 2 .
[0049] where x 1u is the image of the unique part of the visible light image x 1 , which is the image obtained after reconstructing the unique features of the visible light image x 1 . x 2u is the image of the unique part of the near-infrared image x 2 , which is the image obtained after reconstructing the unique features of the near-infrared image x 2 .
[0050] S102. Use the conditional cross-modal entropy model to compress the common features and the unique features of each image to be compressed, and obtain the compressed common features and the compressed unique features of each image to be compressed.
[0051] In the field of image compression, the entropy model refers to a statistical model used to predict and encode the probability distribution of image data. The entropy model is one of the key components in image compression algorithms, used to determine how to most effectively represent image data, thereby achieving a higher compression ratio without significantly losing image quality.
[0052] The conditional cross-modal entropy model is an entropy model used in multi-modal data processing. When compressing image features, the conditional cross-modal entropy model considers the dependence between common features and unique features, and uses the above dependence relationship to optimize encoding or compression to improve encoding efficiency.
[0053] Specifically, the conditional cross-modal entropy model includes multiple cross-modal compression branches, each branch for processing one image feature. One of the cross-modal compression branches is used to compress the common feature to obtain the compressed common feature. The other cross-modal compression branches are respectively used to compress the unique features of each image to be compressed, obtaining the compressed unique features of each image to be compressed. During the process of using one of the cross-modal compression branches to compress the common feature, a latent representation after quantization of the common feature will be obtained. The latent representation after quantization of the common feature is used as an estimation condition and input into the other cross-modal compression branches, so that during the process of the other cross-modal compression branches compressing the unique features of each image to be compressed, the dependency between the common feature and the unique features is considered to improve the encoding efficiency. The number of cross-modal compression branches is the number of images to be compressed plus 1.
[0054] In a possible implementation, as Figure 2 shown, the common feature g C extracted in S101 1 and the unique part of the visible light image x 1u and the unique part of the near-infrared image x 2 are input into the three-branch cross-modal compression module. Through the three-branch cross-modal compression module, the common feature g 2u and the unique part of the visible light image x C and the unique part of the near-infrared image x 1 are compressed to obtain the compressed common feature 1u and the compressed unique part of the visible light image x 2 and the compressed unique part of the near-infrared image x 2u . This process can be represented by formula (4). where, f 1 represents the function of the three-branch cross-modal compression module. The three-branch cross-modal compression module contains three compression branches, and each branch consists of an encoder, an entropy model, and a decoder. For entropy coding, a conditional cross-entropy (CCE) module is proposed. This model uses the latent representation of the quantized common feature as a condition to predict the unique part of the visible light image x 2 and the unique part of the near-infrared image x TCC and the unique part of the near-infrared image x 1 and the unique part of the near-infrared image x 1u
[0055] where, f 2 represents the function of the three-branch cross-modal compression module. The three-branch cross-modal compression module contains three compression branches, and each branch consists of an encoder, an entropy model, and a decoder. For entropy coding, a conditional cross-entropy (CCE) module is proposed. This model uses the latent representation of the quantized common feature as a condition to predict the unique part of the visible light image x 2u and the unique part of the near-infrared image x 1 and the unique part of the near-infrared image x 2Image x of the unique part 2u The probability mass function (PMF). After passing through the TCC module, the compressed common features Are processed through two convolutional layers to obtain the compressed common part And .
[0057] S103. Fuse the compressed common features with the compressed unique features of each image to be compressed respectively to obtain the compressed images corresponding to each image to be compressed.
[0058] For the compressed unique features of each image to be compressed, fuse them with the compressed common features respectively to obtain the compressed images corresponding to each image to be compressed. Specifically, the compressed unique features can be added to the compressed common and unique features respectively to obtain the compressed images corresponding to each image to be compressed.
[0059] In a possible implementation, as Figure 2 Shown, to improve the image reconstruction quality, the compressed common part And , visible light image x 1 Image of the compressed unique part And near-infrared image x 2 Image of the compressed unique part , are further fused through a multi-modal feature fusion module to reconstruct the compressed visible light image And the compressed near-infrared image , as shown in formula (5).
[0060]
[0061] In the embodiments of the present application, the features in each image to be compressed are divided into two parts, namely, the common features shared by multiple images to be compressed and the unique features exclusive to each image to be compressed. The common and unique features are compressed separately and then fused in the feature domain to reconstruct the compressed image. Since the common features are encoded only once, the repeated compression of the common features in multiple images is avoided, saving the coding bit rate and reducing the pressure of storage and transmission.
[0062] On the basis of the above embodiments, the embodiments of the present application further optimize the step of "S101. Extract the common features of multiple images to be compressed and the unique features of each image to be compressed", as Figure 3 Shown, the optimized image feature extraction process mainly includes S201-S202.
[0063] S201. Extract the preliminary common features of multiple images to be compressed.
[0064] The initial common features refer to the features obtained by extracting the common features of multiple images to be compressed for the first time.
[0065] In the embodiment of the present application, the method for extracting the common features of images is mainly performed by Figure 2 the CFD module 21 in the DeepRNC network architecture. Specifically, the CFD module plays an important role in the DeepRNC network, and is mainly used to separate the common features and unique features between multiple images to be compressed.
[0066] In the embodiment of the present application, 2 images to be compressed are taken as an example for illustration. The CFD module is obtained by improving the CUNet. The CFD module represents each image to be compressed with two different convolutional dictionaries, as shown in formula (6).
[0067]
[0068] Among them, x 1 represents the visible light image, x 2 represents the near-infrared image, and * represents the convolution operation. represents the common sparse feature response, represents the unique sparse feature response of the visible light image, represents the unique sparse feature response of the near-infrared image. and are the corresponding common dictionary filters, is the unique dictionary filter of the visible light image, is the unique dictionary filter of the near-infrared image.
[0069] In order to obtain the common feature response and unique feature response of the visible light image and the near-infrared image, the following optimization problem needs to be solved, as shown in formula (7).
[0070]
[0071] Among them, λ c 、λ u and λ v are regularization parameters, and ||·|| 1 represents the l 1 norm.
[0072] In the above formula (7), there are three variables to be optimized. In the embodiment of the present application, a strategy of alternately updating each variable is adopted, that is, when optimizing one variable, the other variables are fixed and alternated. For example, when updating the common feature response c k , the unique feature response u sThe unique feature response v of the visible and near-infrared images t Remain fixed, and there will be and Thus, the common feature response c k Can be updated by solving the following optimization problem, as shown in Equation (8).
[0073]
[0074] Where is and The concatenation of, is and The concatenation of. Equation (8) is a standard convolutional sparse coding problem, which can be solved by the Learned Convolutinoal Sparse Coding (LCSC) algorithm, and the solution obtained is as follows in Equation (9).
[0075]
[0076] Among them, C is the stacking of the common feature responses , C l Is the result after the l-th iterative update of C. Is the learnable convolutional dictionary filter, S λ Is the soft threshold operator. From the above update Of the process, the CFE module can be obtained. As Figure 4 Shown, the specific structure of the CFE module includes: the convolutional dictionary filter E c , the soft threshold operator S λ , L LCSC blocks, the convolutional dictionary filter D c , the convolutional dictionary filter E c , the soft threshold operator S λ .
[0077] The unique feature response u of the visible light image s And the unique feature response v of the near-infrared image t Can be updated in the same way as the common feature update method to obtain the Unique Feature Extraction (UFE) module. The UFE module has a similar structure to the CFE module, except for the input. For example, when updating the unique feature response u s Of the visible light image, the common feature response is fixed. In addition, since the unique feature response v t Of the near-infrared image is assumed to be the same as the unique feature response u sIrrelevant. The unique features of the near-infrared image respond to v t Can also be ignored, so the unique feature response u of the visible light image s Can be updated by solving the following optimization problem, as shown in formula (10).
[0078]
[0079] Where So the previously mentioned LCSC algorithm can be used for solving. However, when solving the optimization problem shown in formula (10), Is an unknown quantity, and Depends on the unique feature response u of the visible light image s . To solve this contradiction, the CU-Net+ network structure is proposed in the embodiments of the present application, that is, the optimization problem is solved in a cyclic manner.
[0080] S202. Extract the common features of multiple images to be compressed based on the multiple images to be compressed and the preliminary common features of the multiple images to be compressed.
[0081] The common feature g output by the CFD module C Is the common feature c extracted in the second stage k 2 , And the image of the common part can also be obtained from the common feature By filtering with the common part dictionary filter, as shown in the following formula (11).
[0082]
[0083] Where And Are the convolution dictionary filters of the common part, sharing weights with f in formula (2) C1 And f in formula (3) C2 .
[0084] Furthermore, as an important preprocessing step, the CFD module will be pre-trained by the loss function shown in formula (12).
[0085]
[0086] Where And Are the visible light image and the near-infrared image output from the CFD module as shown in Figure 5 .
[0087] Referring to the idea of CU-Net+, as Figure 5As shown, in the embodiment of the present application, the CFD module is divided into two stages. In the first stage, the preliminary unique features of the visible light image, the preliminary unique features of the near-infrared image, and the preliminary common features are extracted. In this way, the input of the second stage can be consistent with formula (10), and the common feature g of multiple images to be compressed can be extracted. C .
[0088] In the embodiment of the present application, the CFD module divides the common features shared by the visible light image and the near-infrared image into two modalities and the unique features of each single modality. Therefore, duplicate compression of the common features can be avoided, and the coding bit rate can be significantly saved.
[0089] On the basis of the above embodiment, the embodiment of the present application further optimizes the step of "S102. Using the conditional cross-modal entropy model to compress the common features and the unique features of each of the images to be compressed, to obtain the compressed common features and the compressed unique features of each of the images to be compressed", as Figure 6 shown. The optimized image feature extraction process mainly includes S301-S302.
[0090] The feature compression method provided in the embodiment of the present application is mainly executed by the TCC module 22 in the Figure 2 DeepRNC network architecture in.
[0091] S301. Encode the common features and the unique features of each of the images to be compressed respectively, to obtain the latent representations of the common features and the latent representations of each of the unique features.
[0092] As Figure 2 shown, the TCC module includes 3 compression branches, and each compression branch includes an encoder, an entropy model, and a decoder. The encoder and the decoder in each compression branch follow the design of the attention-based residual model. After passing through the encoder, the latent representation y i (i = 1, 2, 3) corresponding to the features of each compression branch can be obtained.
[0093] Specifically, as Figure 2 shown, the unique part of the visible light image x 1 is input into encoder 1, and encoder 1 outputs the latent representation y 1u of the unique features of the visible light image x 1 ; the unique part of the near-infrared image x 1 is input into encoder 2, and encoder 2 outputs the latent representation y 2 of the unique features of the near-infrared image x 2u ; the unique part of the near-infrared image x 2Potential representation y of unique features 2 , common feature g c In the input encoder 3, the encoder 3 outputs the potential representation y of the common feature 3 .
[0094] S302. Use an autoregressive entropy model to perform entropy coding on the potential representation of the common feature, and obtain the potential representation of the quantized common feature
[0095] For the entropy coding of the common feature, in the embodiment of the present application, a context-based autoregressive entropy model is used to perform entropy coding on the potential representation of the common feature, as shown in formula (13).
[0096]
[0097] In formula (13), is the quantized hyperprior potential representation is the j-th element of denotes the context pixels of 3,j and are respectively the mean and variance of the Gaussian distribution . Then, the bit rate R of the compression branch of the common feature 3 can be obtained through and the likelihood estimation of, as shown in formula (14).
[0098]
[0099] S303. Conditional on the probability distribution of the potential representation of the common feature, use each of the cross-modal conditional entropy models to perform entropy coding on the potential representation of each of the unique features, and obtain the potential representation of each of the quantized unique features
[0100] In a possible implementation, conditional on the probability distribution of the potential representation of the common feature, use each of the cross-modal conditional entropy models to perform entropy coding on the potential representation of each of the unique features, and obtain the potential representation of each of the quantized unique features, including: conditional on the probability distribution of the potential representation of the common feature, calculate the probability distribution of each of the unique features respectively; perform likelihood estimation on the probability distribution of each of the first unique features respectively to obtain the bit stream of each of the unique features; perform arithmetic decoding on the bit stream of each of the unique features respectively to obtain the potential representation of each of the quantized unique features
[0101] For and For the entropy coding, in the embodiments of the present application, a conditional cross-modal entropy (CCE) model is proposed, which considers the dependencies between the common and unique parts. In the CCE model, As an estimate And The conditions for the probability distribution. Similar to It is, the probability distribution of Is modeled as a Gaussian distribution. Considering The j-th element of, each The mean μ i,j And variance Are determined by the quantized hyper-prior latent representation , the context latent representation And the condition Common decision. Thus, The probability distribution of can be expressed as formula (15).
[0102]
[0103] As Figure 7 Shown, through the conditional probability of formula (15), the mean μ i,j And variance Can be predicted by the parameter prediction unit, as shown in formula (16).
[0104]
[0105] In formula (16), ψ i Represents the hyper-prior parameter, f h Is the function of the hyper-prior decoder. φ i j Is the context information predicted by the context prediction unit from , f cp Is the function of the context prediction unit. ψ i And φ i j Are concatenated with As the input of the parameter prediction unit, and its function is represented by f pp . Then, the bit rate of the unique feature compression branch can be obtained from And The likelihood estimate of, as shown in formula (17).
[0106]
[0107] Finally, the estimated total bit rate R can be obtained by adding the bit rates of the three branches, as shown in formula (18).
[0108] R = R 1 + R 2 + R3· (18)
[0109] S304. Decode the potentially represented common features and the potentially represented unique features after quantization of each unique feature respectively to obtain the compressed common features and each compressed unique feature.
[0110] In the embodiments of the present application, a conditional cross-modal entropy model is proposed to fully explore the dependence between the common features and the unique features when compressing features, which helps to improve the encoding efficiency.
[0111] Based on the above embodiments, the embodiments of the present application further optimize the step of "S103. Fuse the compressed common features with the compressed unique features of each image to be compressed respectively to obtain the compressed image corresponding to each image to be compressed", such as Figure 8 shown. The optimized feature fusion process mainly includes S401-S404.
[0112] After obtaining the compressed common part and the unique part from the TCC module, DeepRNC passes through a multi-modal feature fusion (MF) module to reconstruct the visible light image and the near-infrared image from these compressed features. The structure of the MFF module is as Figure 9 shown. This module consists of two fusion branches.
[0113] S401. Add the compressed common features and the compressed unique features of each one respectively to obtain the preliminary fusion image corresponding to each image to be compressed.
[0114] Specifically, the compressed common part image and the unique part image are first added according to the convolutional sparse coding model in formula (19) to reconstruct the preliminary fusion image of the visible light image and the preliminary fusion image of the near-infrared image :
[0115]
[0116] S402. Perform convolutional processing on each preliminary fusion image respectively to obtain the initial feature map corresponding to each preliminary fusion image.
[0117] Specifically, the preliminary fusion image of the visible light image and the preliminary fusion image of first pass through the convolutional layer to obtain p 1 and p 2 as the input of the cross-modal Transformer block.
[0118] S403. Adjust the attention and enhance the features of the initial feature maps corresponding to each preliminary fusion image to obtain each enhanced feature map.
[0119] The preliminarily fused images are further improved in quality through a group of cross-modal Transformer blocks and enhancement blocks.
[0120] In a possible implementation, adjusting the attention and enhancing the features of the initial feature maps corresponding to each preliminary fusion image to obtain each enhanced feature map includes: performing cross-attention operations on each initial feature map to obtain cross feature maps; respectively performing addition operations, multi-layer perception operations, and feature enhancement operations on each initial feature map and the cross feature maps to obtain each enhanced feature map.
[0121] Specifically, and First, p 1 and p 2 are obtained through a convolutional layer as the inputs to the cross-modal Transformer block. As Figure 9 shown, taking the right fusion branch as an example, the inputs to the cross-modal Transformer block are p 1 and p 2 , and the output is obtained through the following formula (20).
[0122]
[0123] where Attn(·,·) represents the cross-modal attention operation, LN(·) represents layer normalization, and MLP represents the multi-layer perceptron. The cross-attention operation can be expressed as formula (21).
[0124]
[0125] In formula (20), . Q 1 = W Q X 1 , K 2 = W K X 2 , are respectively the query, key, and value of the attention mechanism. W Q , W K is a learnable weight matrix. B is the bias parameter, and d is a learnable parameter that controls the size of the softmax input. Note that the attention mechanism here is channel-wise. In the MFF module, the cross-modal Transformer blocks in the left and right fusion branches have the same structure, but the inputs are swapped. After passing through the cross Transformer blocks, the output is further input into an enhancement block for feature enhancement. As Figure 9 shown, the enhancement block is relatively simple and consists of only three residual blocks. After passing through the enhancement block, we can obtain the enhanced features as shown in Equation (22).
[0126]
[0127] where f EN represents the function of the enhancement block. Note that there are a total of n cascaded cross-modal Transformers and enhancement blocks, and the outputs of the last enhancement block are denoted as and .
[0128] S404. Image reconstruction is performed separately based on each enhanced feature map to obtain the compressed images corresponding to each image to be compressed.
[0129] The finally reconstructed images and are obtained by inputting and into the reconstruction convolutional layer, as shown in Equation (23).
[0130]
[0131] The multi-modal feature fusion (MFF) module provided in the embodiments of the present application fuses compressed features through cross Transformers and enhancement blocks.
[0132] Furthermore, the compression performance of the DeepRNC framework proposed in the embodiments of the present application was evaluated on three RGB-NIR datasets, and the experimental results show that higher compression efficiency can be achieved compared to existing single-modal or multi-modal image compression methods.
[0133] To improve the multi-modal compression performance, the DeepRNC network needs to adopt a multi-stage training method. In the first stage, the CFD module is pre-trained using the loss function in Equation (12) . Subsequently, in the second stage, the pre-trained CFD module is loaded, and the CFD and TCC modules are jointly trained. The loss function used is the following rate-distortion loss function , such as formula (24).
[0134]
[0135] Where R represents the total bit rate estimated by formula (18); D represents the image distortion measured by the mean square error (MSE). are the visible light image and the near-infrared image after rough fusion of the formula. In the third stage, fix the parameters of the CFD and TCC modules, and train the MFF module alone. The loss function is only the image distortion loss such as formula (25).
[0136]
[0137] Where and are the visible light image and the near-infrared image finally reconstructed by the MFF module. Finally, fine-tune the entire DeepRNC network end-to-end, and the loss function adopted is the rate-distortion loss function such as formula (26).
[0138]
[0139] Figure 10 is a schematic structural diagram of an image compression device in an embodiment of the present application, as Figure 10 shown, the image compression device 100 provided by the embodiment of the present application mainly includes:
[0140] A feature extraction module 101, configured to extract the common features and the unique features of multiple images to be compressed, where the multiple images to be compressed are images obtained in different bands for the same scene; a feature compression module 102, configured to use a conditional cross-modal entropy model to compress the common features and the unique features of each image to be compressed, so as to obtain the compressed common features and the compressed unique features of each image to be compressed; a feature fusion module 103, configured to fuse the compressed common features with the compressed unique features of each image to be compressed respectively, so as to obtain the compressed images corresponding to each image to be compressed.
[0141] In a possible implementation manner, the feature extraction module 101 is specifically configured to extract the preliminary common features of multiple images to be compressed; based on the multiple images to be compressed and the preliminary common features of the multiple images to be compressed, extract the common features of the multiple images to be compressed.
[0142] In a possible implementation, the conditional cross-modal entropy model includes an autoregressive entropy model and multiple cross-modal conditional entropy models; the feature compression module 102 is specifically configured to encode the common features and the unique features of each image to be compressed respectively, so as to obtain the latent representation of the common features and the latent representation of each unique feature; use the autoregressive entropy model to perform entropy encoding on the latent representation of the common features to obtain the quantized latent representation of the common features; use each cross-modal conditional entropy model to perform entropy encoding on the latent representation of each unique feature conditional on the probability distribution of the latent representation of the common features, so as to obtain the quantized latent representation of each unique feature; decode the quantized latent representation of the common features and the quantized latent representation of each unique feature respectively to obtain the compressed common features and the compressed unique features of each image.
[0143] In a possible implementation, the feature compression module 102 is specifically configured to calculate the probability distribution of each unique feature respectively conditional on the probability distribution of the latent representation of the common features; perform likelihood estimation on the probability distribution of each first unique feature respectively to obtain the bitstream of each unique feature; perform arithmetic decoding on the bitstream of each unique feature respectively to obtain the quantized latent representation of each unique feature.
[0144] In a possible implementation, the feature fusion module 103 is specifically configured to add the compressed common features and the compressed unique features of each image respectively to obtain the preliminary fusion image corresponding to each image to be compressed; perform convolution processing on each preliminary fusion image respectively to obtain the initial feature map corresponding to each preliminary fusion image; perform attention adjustment and feature enhancement on the initial feature map corresponding to each preliminary fusion image to obtain each enhanced feature map; perform image reconstruction based on each enhanced feature map respectively to obtain the compressed image corresponding to each image to be compressed.
[0145] In a possible implementation, the feature fusion module 103 is specifically configured to perform cross-attention operation on each initial feature map to obtain a cross feature map; perform addition operation, multi-layer perception operation and feature enhancement operation on each initial feature map and the cross feature map respectively to obtain each enhanced feature map.
[0146] In a possible implementation, the multiple images to be compressed include visible light images and infrared images.
[0147] The image compression device provided by the embodiments of the present application can execute the image compression method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0148] Figure 11 It is a schematic structural diagram of an electronic device provided in this embodiment. The electronic device may include an image compression device, such as Figure 11As shown, the electronic device 1100 includes a processor 1110, a memory 1120, an input device 1130, and an output device 1140; the number of processors 1110 in the electronic device can be one or more, Figure 11 here, one processor 1110 is taken as an example; the processor 1110, the memory 1120, the input device 1130, and the output device 1140 in the electronic device can be connected through a bus or other means, Figure 11 here, being connected through a bus is taken as an example.
[0149] The memory 1120, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the image compression method in the embodiments of the present invention. The processor 1110 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the memory 1120, that is, implements the image compression method provided by the embodiments of the present invention.
[0150] The memory 1120 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory 1120 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 1120 may further include a memory remotely set relative to the processor 1110, and these remote memories can be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0151] The input device 1130 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the electronic device, and may include a keyboard, a mouse, etc. The output device 1140 may include a display device such as a display screen.
[0152] This embodiment also provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to implement the image compression method provided by the embodiments of the present invention when executed by a computer processor.
[0153] Of course, for a storage medium containing computer-executable instructions provided by the embodiments of the present invention, the computer-executable instructions are not limited to the method operations as described above, and can also execute related operations in the image compression method provided by any embodiment of the present invention.
[0154] From the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software and necessary general-purpose hardware. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk, or optical disc of a computer, etc., and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0155] It should be noted that in the embodiments of the above image compression device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.
[0156] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including an..." does not exclude the presence of additional identical elements in the process, method, article or device including the said element.
[0157] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments described herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. An image compression method, characterized in that: The method comprises: Extracting common features of a plurality of images to be compressed and unique features of each of the images to be compressed, wherein the plurality of images to be compressed are images acquired in different bands for the same scene; Compressing the common features and the unique features of each of the images to be compressed using a conditional cross modal entropy model to obtain compressed common features and compressed unique features of each of the images to be compressed; The compressed common features are respectively fused with the compressed unique features of each of the images to be compressed to obtain compressed images corresponding to each of the images to be compressed.
2. The method according to claim 1, characterized in that: Extract the common features of multiple images to be compressed, including: Extracting preliminary common features of multiple images to be compressed; Based on a plurality of images to be compressed and preliminary common features of the plurality of images to be compressed, common features of the plurality of images to be compressed are extracted.
3. The method according to claim 2, characterized in that The conditional cross-modal entropy model includes an autoregressive entropy model and a plurality of cross-modal conditional entropy models; The common features and the unique features of each of the images to be compressed are compressed using a conditional cross modal entropy model to obtain the compressed common features and the compressed unique features of each of the images to be compressed, including: Encoding the common features and the unique features of each of the images to be compressed respectively to obtain a potential representation of the common features and a potential representation of each of the unique features; Performing entropy coding on the potential representation of the common features by using the autoregressive entropy model to obtain a quantized potential representation of the common features; Taking the probability distribution of the potential representation of the common features as a condition, using each of the cross-modal conditional entropy models, entropy encoding the potential representation of each of the unique features to obtain a quantized potential representation of each of the unique features; The quantized potential representations of the common features and the quantized potential representations of the unique features are decoded respectively to obtain the compressed common features and the compressed unique features.
4. The method according to claim 3, characterized in that: Taking the probability distribution of the potential representation of the common features as a condition, using each of the cross-modal conditional entropy models, entropy encoding the potential representation of each of the unique features to obtain a quantized potential representation of each of the unique features includes: Calculating the probability distribution of each of the unique features respectively based on the probability distribution of the potential representation of the common features; Performing likelihood estimation on the probability distribution of each of the first unique features respectively to obtain a bit stream of each of the unique features; The bit streams of the respective unique features are respectively arithmetically decoded to obtain quantized potential representations of the respective unique features.
5. The method according to claim 3, characterized in that: The compressed common features are respectively fused with the compressed unique features of each of the images to be compressed to obtain compressed images corresponding to each of the images to be compressed, including: The compressed common features are added to the compressed unique features to obtain preliminary fused images corresponding to the images to be compressed; Perform convolution processing on each of the preliminary fused images to obtain an initial feature map corresponding to each of the preliminary fused images; Performing attention adjustment and feature enhancement on the initial feature maps corresponding to each of the preliminary fused images to obtain each enhanced feature map; Image reconstruction is performed based on each of the enhanced feature maps to obtain a compressed image corresponding to each image to be compressed.
6. The method according to claim 5, characterized in that The step of performing attention adjustment and feature enhancement on the initial feature maps corresponding to each of the preliminary fused images to obtain each enhanced feature map includes: Performing a cross attention operation on each of the initial feature maps to obtain a cross feature map; Each of the initial feature maps and the cross feature map is respectively subjected to an addition operation, a multi-layer perception operation and a feature enhancement operation to obtain each enhanced feature map.
7. The method according to any one of claims 1 to 6, characterized in that: The multiple images to be compressed include visible light images and infrared images.
8. An image compression device, characterized in that: The device comprises: A feature extraction module, used to extract common features of multiple images to be compressed and unique features of each of the images to be compressed, wherein the multiple images to be compressed are images acquired in different bands for the same scene; A feature compression module, used for compressing the common features and the unique features of each of the images to be compressed using a conditional cross-modal entropy model to obtain compressed common features and compressed unique features of each of the images to be compressed; The feature fusion module is used to fuse the compressed common features with the compressed unique features of each image to be compressed to obtain a compressed image corresponding to each image to be compressed.
9. An electronic device, characterized in that: The electronic device comprises: one or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the image compression method according to any one of claims 1 to 7.
10. A storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the image compression method according to any one of claims 1 to 7 is implemented.