Characterization domain sampling and mixed Transform image compressed sensing method
By using characterization domain sampling and hybrid Transformer methods in image compression perception, cross-domain feature alignment, multi-scale information fusion and model volume expansion are solved, and efficient and lightweight image reconstruction is achieved, with excellent reconstruction performance and robustness.
Patent Information
- Application Number
- CN202510171839.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-17
AI Technical Summary
Existing image compression perception methods have challenges in cross-domain feature alignment, multi-scale information fusion, and model volume expansion, resulting in poor reconstruction quality and computational efficiency.
Using the method of characterizing domain sampling and hybrid Transformer, deep representation is learned through multi-layer convolution and nonlinear activation functions, combining the depth gradient descent module and the Transformer module to perform fine reconstruction and information fusion, replacing the traditional near-end projection term.
It achieves high reconstruction performance, detail retention and texture reconstruction, and can maintain excellent performance at low sampling rates, with strong robustness and versatility.
Smart Images

Figure CN119991835A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image compression sensing, and in particular to an image compression sensing method of representation domain sampling and hybrid Transformer. Background Art
[0002] Compressed sensing (CS) is a signal processing theory that breaks through the limitations of the traditional Shannon-Nyquist sampling theorem. It can significantly reduce sampling costs and data storage requirements by recovering sparse or compressible signals from sampled data at significantly lower sampling rates than traditional sampling rates. Therefore, CS theory has been widely used in many fields, including snapshot compression imaging, medical magnetic resonance imaging (MRI), hyperspectral compression imaging, and laser scanning imaging.
[0003] The core issues of CS methods mainly include sampling strategies and reconstruction strategies. At present, a variety of sampling matrices, such as random matrices, binary matrices, and structural matrices, have been proposed to solve the sampling strategy problem. In terms of reconstruction strategy, methods such as convex optimization algorithms, greedy algorithms, and iterative threshold algorithms are widely used. However, these methods have some problems, such as slow convergence and limited modeling capabilities for complex signals. These algorithms not only constrain the consistency of reconstruction and measurement values, but also usually encourage the sparsity of solutions. Prior terms usually involve sparse operators related to predefined transformation bases (such as L1 regularization of ISTA, discrete cosine transform (DCT)). These methods perform well in convergence and mathematical analysis, but face problems of high computational complexity and poor adaptability. In recent years, the outstanding capabilities of deep neural networks have spawned a variety of deep compressed sensing algorithms (DCS). One type of method adopts a purely data-driven CS architecture and directly trains the model to learn the potential inverse mapping from a large amount of data. Although these methods can automatically solve CS problems, their black-box characteristics are accompanied by a large number of parameters and computational trial and error, which is very inefficient. In contrast, the Deep Unfolded Networks (DUNs) strategy advocates the complementarity of neural networks and optimization algorithms, replacing the manual calculation of optimization terms / prior terms with the learnable parameters of neural networks (CNN / Transformer). Therefore, this DCS method takes into account both reconstruction quality and mathematical interpretability.
[0004] However, the current DCS methods still have three key challenges: Cross-domain feature alignment loss: compression is performed in the low-level pixel domain, but reconstruction is based on the high-level feature domain. Frequent cross-domain gradient descent leads to significant feature alignment loss. Insufficient multi-scale information fusion: Feature expressions fail to fully integrate multi-scale information. They usually use traditional proximal projection terms and fail to exert the fitting learning ability of prior terms. Model volume expansion: The DCS model exchanges a small improvement in quality at the cost of exponential expansion of volume, which is easily limited by end-side equipment (such as lasers, MRI, etc.).
[0005] Therefore, how to overcome these challenges and design an efficient, lightweight compressed sensing method with high reconstruction quality has become an important direction of current research. Summary of the invention
[0006] The purpose of this section is to summarize some aspects of embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the specification abstract and the invention title of this application to avoid blurring the purpose of this section, the specification abstract and the invention title, and such simplifications or omissions cannot be used to limit the scope of the present invention.
[0007] In view of the problems existing in the above-mentioned existing representation domain sampling and hybrid Transformer image compression sensing methods, the present invention is proposed.
[0008] Therefore, the purpose of the present invention is to provide an image compression sensing method of representation domain sampling and hybrid Transformer, which innovatively designs a representation compression sensing (RCS) sub-model. Different from other direct sampling CS methods, RCS advocates extracting compact feature representations with high information density and rich semantics in the representation space first, so as to complete a more efficient signal compression and reconstruction process in the representation domain.
[0009] In order to solve the above technical problems, the present invention provides the following technical solutions: an image compression sensing method of representation domain sampling and hybrid Transformer, comprising the following steps:
[0010] Sampling recovery stage: learning the input image through multiple layers of progressive 5×5 and 7×7 convolutions and nonlinear activation functions The deep representation of the fragmented window information is fully interactive, and then the aggregation features containing both low and high levels are generated through jump connections. And on this basis, the downsampled Y is obtained;
[0011] Fine reconstruction iteration stage: Use the deep gradient descent module to perform optimization updates of the reconstruction features. The deep gradient descent module expands the calculation of the update fidelity term to the neural network:
[0012]
[0013] Among them, the nonlinear network f res (·) improves the ability to extract residual information from observations, while and As the projection mapping of the observation domain and the depth feature domain respectively; Recorded as the intermediate result of the recovery process; As a pseudo-inverse mapping of the sample S;
[0014] The proximal projection is replaced by a three-scale sparse denoising sub-network; the Transformer module is then used to perform fine complementary fusion of the gradient descent terms. The Transformer module includes a cross-attention module, a windowed local attention module, and a FeedForwardNetwork module.
[0015] As a preferred solution of the image compression sensing method of the representation domain sampling and hybrid Transformer of the present invention, wherein: in the sampling, It is divided into non-overlapping blocks of size P×P (the P value is set to 32 in the experiment) and then processed by unbiased linear convolution. Sampling to get the compressed value Where r is the sampling rate, the image length and width are H and W, the size of the convolution kernel Φ is P×P, the stride is P, and the sampling subnetwork S(·) is expressed as:
[0016]
[0017] Encoder is a feature encoder; deconvolution with shared parameters with the sampling kernel Φ is applied to the representation compression result Y. Get the initial representation estimate
[0018] The compressed and restored representation is transformed back to the original image domain, and the initial recovery subnetwork It is expressed as:
[0019]
[0020] Reconstructing subnetworks is the initial pseudo-inverse of the sampling subnetwork S(·); Decoder is the feature decoder.
[0021] As a preferred solution of the image compression sensing method of the representation domain sampling and hybrid Transformer of the present invention, the deep gradient descent module first calibrates the deep reconstruction features at each stage through the observation domain residual information to ensure the optimization of the features. The consistency with the observed value Y.
[0022] As a preferred solution of the image compression sensing method of the representation domain sampling and hybrid Transformer of the present invention, wherein: the three-scale sparse denoising performs three downsampling of the features by 2×2 convolution with a step size of 2 and GELU (a specific activation function), keeps the number of channels unchanged, and gradually reduces the size of the feature map to Generate feature representations of different scales. Subsequently, use linear interpolation to upsample the low-resolution features by ×2 and concatenate them with the features of the corresponding resolution through residual connection. Use 5×5 depth convolution to aggregate feature information with a larger window.
[0023] As a preferred solution of the image compression sensing method of the representation domain sampling and hybrid Transformer of the present invention, the auxiliary feature is introduced in the three-scale sparse denoising process. The residual block combines the optimized features with the original X (k) The layers are pooled along the channel dimension and output through a 3×3 convolution.
[0024] As a preferred solution of the image compression sensing method of the representation domain sampling and hybrid Transformer of the present invention, the specific input Q of the cross attention module comes from K and V, which are projected to The new component, after the Softmax operation (a probabilistic activation function), generates a transposed attention map of the cross-CS level features Cross-Attention Module The calculation of is defined as:
[0025]
[0026] Conv 1×1 Denoted as linear convolution; in the specific iteration of the network, the pseudo-inverse of the measured feature is As the Q1 query component, the cross attention calculation is introduced; the K1 and V1 key values are both from the reconstruction result component of the current stage
[0027]
[0028] Through the dual-domain feature cross calculation It is considered as a potential error term of the PGD (proximal gradient descent strategy) fidelity constraint term, and then it is re-input into the second stage GCA as K2 and V2 components, which is consistent with the current As the new Q2, a second cross-attention calculation is performed. It is a refined supplementary item of the first stage of PGD after modeling the global information, which is used to guide Further updates:
[0029]
[0030] As a preferred solution of the image compression sensing method of the representation domain sampling and hybrid Transformer of the present invention, the FeedForwardNetwork module is composed of LayerNorm layer, linear layer and DConv3×3 (Deep Grouped Convolution) composition.
[0031] As a preferred solution of the image compression sensing method of the representation domain sampling and hybrid Transformer of the present invention, the window local attention module uses a sliding window of deep convolution to generate a dynamic local attention map A′, and aggregates local features through channel-by-channel convolution, specifically including:
[0032]
[0033] Among them, f σ is the GeLU activation function, which is used to produce a smoother attention distribution and enhance the input The local information representation capability of , W' is the projection matrix of the corresponding component; then, the output of the window local attention module is residually connected with the original features and further processed through a customized feedforward network:
[0034]
[0035] In the loss function, RHT-Net defines the error of the image pair as And adopt hybrid loss function for end-to-end optimization:
[0036]
[0037] Where N is the number of training sample pairs, Θ represents all trainable parameters; the ratios of λ′ and λ″ are both set to 0.5 to balance the robustness and convergence speed of the model.
[0038] Beneficial effects of the present invention:
[0039] 1. High reconstruction performance: The method of the present invention (RHT-Net) achieves better reconstruction quality than competing methods at various sampling rates. Especially at low sampling rates, the model of the present invention can still maintain excellent performance, and the reconstructed image has higher clarity and accuracy.
[0040] 2. Detail preservation and texture reconstruction: Compared with other methods, the model of the present invention performs well in processing detail and texture features. Through the powerful feature extraction and global modeling capabilities of the representation domain sampling (RCS) module and the refined reconstruction module, the RHT-Net model can better restore the detail features in the image and provide clearer edge and texture information.
[0041] 3. Robustness and versatility: The RHT-Net model proposed in this paper has been fully tested on multiple benchmark datasets and has achieved better reconstruction quality at different sampling rates. After Gaussian noise is added to the image, it is still ahead of other methods. This shows that the method has strong robustness and versatility and is applicable to various sampling rates and image scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor. Among them:
[0043] Figure 1 :The overall design framework diagram of the network model (RHT-Net) of the present invention;
[0044] Figure 2 : Schematic diagram of the characterization compressed sensing module (RCS) included in the present invention;
[0045] Figure 3 : Schematic diagram of the deep fine reconstruction stage included in the present invention;
[0046] Figure 4 : A quantitative comparison chart of the model proposed in this invention with other advanced methods in public data;
[0047] Figure 5 : Visualization diagram of the proposed model and other advanced methods in public data;
[0048] Figure 6 : Computational cost and performance comparison diagram of the model proposed in the present invention;
[0049] Figure 7 : Schematic diagram of the noise fading robustness of the model proposed in the present invention. DETAILED DESCRIPTION
[0050] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.
[0051] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0052] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.
[0053] Secondly, the present invention is described in detail with reference to the schematic diagram. When describing the embodiments of the present invention in detail, for the sake of convenience, the cross-sectional diagrams showing the device structure will not be partially enlarged according to the general scale, and the schematic diagrams are only examples, which should not limit the scope of protection of the present invention. In addition, in actual production, the three-dimensional dimensions of length, width and depth should be included.
[0054] Reference Figure 1 - Figure 7 , provides an image compression sensing method for representation domain sampling and hybrid Transformer, including the following steps:
[0055] Sampling recovery stage: learning the input image through multiple layers of progressive 5×5 and 7×7 convolutions and nonlinear activation functions The deep representation of the fragmented window information is fully interactive, and then the aggregation features containing both low and high levels are generated through jump connections. And on this basis, the downsampled Y is obtained;
[0056] Fine reconstruction iteration stage: Use the deep gradient descent module to perform optimization updates of the reconstruction features. The deep gradient descent module expands the calculation of the update fidelity term to the neural network:
[0057]
[0058] Among them, the nonlinear network f res (·) improves the ability to extract residual information from observations, while and As the projection mapping of the observation domain and the depth feature domain respectively; Recorded as the intermediate result of the recovery process; As a pseudo-inverse mapping of the sample S;
[0059] Replace the proximal projection with a three-scale sparse denoising sub-network;
[0060] The Transformer module is then used to perform a fine-grained complementary fusion of the gradient descent terms. The Transformer module includes a cross-attention module, a windowed local attention module, and a FeedForwardNetwork module.
[0061] Among them, in the sampling, It is divided into non-overlapping blocks of size P×P (the P value is set to 32 in the experiment) and then processed by unbiased linear convolution. Sampling to get the compressed value Where r is the sampling rate, the image length and width are H and W, the size of the convolution kernel Φ is P×P, the stride is P, and the sampling subnetwork S(·) is expressed as:
[0062]
[0063] Encoder is a feature encoder; deconvolution with shared parameters with the sampling kernel Φ is applied to the representation compression result Y. Get the initial representation estimate
[0064] The compressed and restored representation is transformed back to the original image domain, and the initial recovery subnetwork It is expressed as:
[0065]
[0066] Decoder is the feature decoder; reconstruction sub-network is the initial pseudo-inverse of the sampling subnetwork S(·).
[0067] Furthermore, the deep gradient descent module first calibrates the deep reconstruction features at each stage through the observation domain residual information to ensure the optimization of features The consistency with the observed value Y.
[0068] Among them, the three-scale sparse denoising downsamples the features three times through 2×2 convolution with a step size of 2 and GELU activation, keeping the number of channels unchanged, and gradually reducing the size of the feature map to Generate feature representations of different scales. Subsequently, use linear interpolation to upsample the low-resolution features by ×2 and concatenate them with the features of the corresponding resolution through residual connection. Use 5×5 depth convolution to aggregate feature information with a larger window.
[0069] Specifically, auxiliary features are introduced in the three-scale sparse denoising process The residual block combines the optimized features with the original X (k) The layers are pooled along the channel dimension and output through a 3×3 convolution.
[0070] Furthermore, the specific input Q of the cross attention module comes from K and V, which are projected to The new component, after the Softmax operation (a probabilistic activation function), generates a transposed attention map of the cross-CS level features Cross-Attention Module The calculation of is defined as:
[0071]
[0072] Conv 1×1 Denoted as linear convolution; in the specific iteration of the network, the pseudo-inverse of the measured feature is As the Q1 query component, the cross attention calculation is introduced; the K1 and V1 key values are both from the reconstruction result component of the current stage
[0073]
[0074] Through the dual-domain feature cross calculation It is considered as a potential error term of the PGD (proximal gradient descent strategy) fidelity constraint term, and then it is re-input into the second stage GCA as K2 and V2 components, which is consistent with the current As the new Q2, a second cross-attention calculation is performed. It is a refined supplementary item of the first stage of PGD after modeling the global information, which is used to guide Further updates:
[0075]
[0076] Among them, the FeedForwardNetwork module consists of LayerNorm layer, linear layer and DConv 3×3 (deep grouped convolution); the windowed local attention module uses the sliding window of deep convolution to generate a dynamic local attention map A′ and aggregates local features through channel-by-channel convolution, specifically including:
[0077]
[0078] Among them, f σ is the GeLU activation function, which is used to produce a smoother attention distribution and enhance the input The local information representation capability of , W' is the projection matrix of the corresponding component; then, the output of the window local attention module is residually connected with the original features and further processed through a customized feedforward network:
[0079]
[0080] In the loss function, RHT-Net defines the error of the image pair as And adopt hybrid loss function for end-to-end optimization:
[0081]
[0082] Where N is the number of training sample pairs, Θ represents all trainable parameters; the ratios of λ′ and λ″ are both set to 0.5 to balance the robustness and convergence speed of the model.
[0083] Example:
[0084] Model input: In the embodiment of the present invention, the input part uses the BSD500 dataset as the training set, which contains 400 images. The training data is randomly cropped to 96×96 size and randomly flipped to obtain 80,000 sub-images. In order to improve computational efficiency and model robustness, the image is converted to YCbCr color space, and only the Y channel is used for training and testing. The test datasets include Set5, Set11, BSDS100 and Urban100, and the test images are center cropped and resized to 128×128.
[0085] Training details: The Adam optimizer is used to update the model parameters, and the momentum and weight decay are set to 0.9 and 0.999 respectively. The initial learning rate is set to 1e-3 and decayed at the 60th, 90th, 120th, 150th and 180th epochs.
[0086] Training method: The model is trained for 200 epochs using a batch size of 6. The learning rate decay strategy is to multiply the learning rate by 0.3 at the 60th, 90th, 120th, 150th, and 180th epochs.
[0087] Test details: In the test phase, test images with different sampling rates are reconstructed, and the reconstruction quality is evaluated using PSNR and SSIM.
[0088] Experimental platform: The experimental platform uses the PyTorch framework to implement the model and is trained on the Nvidia RTX 3090 GPU.
[0089] High reconstruction performance: The proposed method (RHT-Net) achieves better reconstruction quality than competing methods at various sampling rates. Especially at low sampling rates, the proposed model can still maintain excellent performance, and the reconstructed image has higher clarity and accuracy.
[0090] Detail preservation and texture reconstruction: Compared with other methods, the model of the present invention performs well in processing detail and texture features. Through the powerful feature extraction and global modeling capabilities of the representation domain sampling (RCS) module and the refined reconstruction module, the RHT-Net model can better restore the detail features in the image and provide clearer edge and texture information.
[0091] Robustness and versatility: The proposed RHT-Net model has been fully tested on multiple benchmark datasets and has achieved better reconstruction quality at different sampling rates. After Gaussian noise is added to the image, it is still ahead of other methods. This shows that the method has strong robustness and versatility and is applicable to various sampling rates and image scenarios.
[0092] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A method for image compression sensing using representation domain sampling and hybrid Transformer, characterized in that: The following steps are involved: Sampling recovery stage: learning the input image through multiple layers of progressive 5×5 and 7×7 convolutions and nonlinear activation functions The deep representation of the fragmented window information is fully interactive, and then the aggregation features containing both low and high levels are generated through jump connections. And on this basis, the downsampled Y is obtained; Fine reconstruction iteration stage: Use the deep gradient descent module to perform optimization updates of the reconstruction features. The deep gradient descent module expands the calculation of the update fidelity term to the neural network: Among them, the nonlinear network f res (·) improves the ability to extract residual information from observations, while and As the projection mapping of the observation domain and the depth feature domain respectively; Recorded as the intermediate result of the recovery process; As a pseudo-inverse mapping of the sample S; Replace the proximal projection with a three-scale sparse denoising sub-network; The Transformer module is then used to perform a fine complementary fusion of the gradient descent terms. The Transformer module includes a cross-attention module, a windowed local attention module, and a Feed Forward Network module.
2. The image compression sensing method of representation domain sampling and hybrid Transformer according to claim 1, characterized in that: In sampling, It is divided into non-overlapping blocks of size P×P (the P value is set to 32 in the experiment) and then processed by unbiased linear convolution. Sampling to get the compressed value Where r is the sampling rate, the image length and width are H and W, the size of the convolution kernel Φ is P×P, the stride is P, and the sampling subnetwork S(·) is expressed as: Encoder represents a feature encoder; Apply deconvolution with shared parameters with the sampling kernel Φ to the representation compression result Y Get the initial representation estimate The compressed and restored representation is transformed back to the original image domain, and the initial recovery subnetwork It is expressed as: Reconstructing subnetworks is the initial pseudo-inverse of the sampling subnetwork S(·); Decoder is the feature decoder.
3. The image compression sensing method of representation domain sampling and hybrid Transformer according to claim 1, characterized in that: The deep gradient descent module first calibrates the deep reconstruction features at each stage by observing the residual information in the domain to ensure the optimization of the features. The consistency with the observed value Y.
4. The image compression sensing method of representation domain sampling and hybrid Transformer according to claim 1, characterized in that: The three-scale sparse denoising downsamples the features three times by using 2×2 convolution with a step size of 2 and GELU (a specific activation function), keeping the number of channels unchanged and gradually reducing the size of the feature map to Generate feature representations of different scales. Subsequently, use linear interpolation to upsample the low-resolution features by ×2 and concatenate them with the features of the corresponding resolution through residual connection. Use 5×5 depth convolution to aggregate feature information with a larger window.
5. The image compression sensing method of representation domain sampling and hybrid Transformer according to claim 4 is characterized in that: Introducing auxiliary features in three-scale sparse denoising The residual block combines the optimized features with the original X (k) The layers are pooled along the channel dimension and output through a 3×3 convolution.
6. The image compression sensing method of representation domain sampling and hybrid Transformer according to claim 1, characterized in that: The specific input Q of the cross attention module comes from K and V, which are projected to The new component, after the Softmax operation (a probabilistic activation function), generates a transposed attention map of the cross-level features Cross-Attention Module The calculation of is defined as: Conv 1×1 Denoted as linear convolution; in the specific iteration of the network, the pseudo-inverse of the measured feature is As the Q1 query component, the cross attention calculation is introduced; the K1 and V1 key values are both from the reconstruction result component of the current stage Through the dual-domain feature cross calculation It is considered as a potential error term of the PGD (proximal gradient descent strategy) fidelity constraint term, and then it is re-input into the second stage GCA as K2 and V2 components, which is consistent with the current As the new Q2, a second cross-attention calculation is performed. It is a refined supplementary item of the first stage of PGD after modeling the global information, which is used to guide Further updates:
7. The image compression sensing method of representation domain sampling and hybrid Transformer according to claim 6 is characterized in that: The Feed Forward Network module consists of Layer Norm layer, linear layer and DConv 3×3 (Deep Grouped Convolution) composition.
8. The image compression sensing method of representation domain sampling and hybrid Transformer according to claim 7, characterized in that: The windowed local attention module uses a sliding window of deep convolution to generate a dynamic local attention map A′ and aggregates local features through channel-by-channel convolution, specifically including: Among them, f σ is the GeLU activation function, which is used to produce a smoother attention distribution and enhance the input The local information representation capability of , W' is the projection matrix of the corresponding component; then, the output of the window local attention module is residually connected with the original features and further processed through a customized feedforward network: In the loss function, RHT-Net defines the error of the image pair as And adopt hybrid loss function for end-to-end optimization: Where N is the number of training sample pairs, Θ represents all trainable parameters; the ratios of λ′ and λ″ are both set to 0.5 to balance the robustness and convergence speed of the model.
Citation Information
Patent Citations
Through-the-wall radar human body behavior feature enhancement method and device
CN115240040A
Image compressed sensing reconstruction method and system based on hybrid Transform and CNN (Convolutional Neural Network)
CN117808907A
Compressed sensing iterative reconstruction method combined with image hierarchical features
CN119273782A
Method for image motion deblurring, apparatus, electronic device and medium therefor
US20240404025A1