CNN and Mama-based double-branch depth compressed sensing reconstruction method and system
By combining the state space model and the dual-branch architecture of CNN branches, the effective fusion of global-local features is achieved, solving the balance of computing efficiency and reconstruction quality in compressed-sensing image reconstruction, and improving the image reconstruction effect.
Patent Information
- Application Number
- CN202510720395.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-08-12
AI Technical Summary
The existing compressed-sensing image reconstruction methods are difficult to balance between computing efficiency and reconstruction quality. The local receptive field of convolutional neural networks limits the modeling ability of global dependencies, and the computing complexity and memory consumption of deep learning methods are high, resulting in limited practical applications.
Using a dual-branch depth compression sensing reconstruction method based on CNN and Mamba, combining the state space model and CNN branches, through a multi-scale feature fusion mechanism, the global dependence features of the image are captured and local details are extracted. The sequence modeling branches SATM and CNN feature extraction branches are designed to achieve effective fusion of global-local features.
It significantly improves the quality of image reconstruction, and can maintain global context and multi-scale local features during the conversion process of low-dimensional measurements to high-resolution reconstruction images, adapt to different scenario needs, and improves the reconstruction effect.
Smart Images

Figure CN120472031A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image reconstruction technology, and in particular to a dual-branch deep compressed sensing reconstruction method and system based on CNN and Mamba. Background Art
[0002] With the rapid development of sensing technology and digital devices, the demand for acquiring and processing images, as an important carrier of visual information, continues to grow. However, the massive amount of data generated by high-resolution images poses a severe challenge to storage and computing resources, especially in application scenarios with limited storage and computing capabilities. The theory of compression provides important theoretical support for this problem. This theory shows that when a signal is sparse in a specific transform domain, the original signal can be accurately reconstructed using measurement data that is much lower than the Nyquist sampling rate. However, how to strike a balance between compression rate and reconstruction quality remains an important research issue. Traditional compressed sensing image reconstruction methods rely primarily on iterative optimization algorithms, including basis pursuit and iterative soft thresholding. However, their high computational cost and complex optimization process limit their practical application. The rapid development of deep learning technology has brought new solutions to compressed sensing image reconstruction. Through end-to-end training, deep learning models can directly learn the mapping from low-dimensional measurements to high-dimensional reconstructed images from data, significantly improving reconstruction efficiency and effectiveness. However, existing deep learning methods still face some key challenges. Convolutional neural networks exhibit good computational efficiency during reconstruction, but their local receptive field limits their ability to model global dependencies, making them particularly ineffective when processing complex image structures. To address these limitations of CNNs, researchers have attempted to introduce the Transformer architecture, which uses a self-attention mechanism to capture global features. However, the Transformer's high computational complexity and memory consumption pose significant challenges for practical applications. Therefore, in order to solve the above problems, a new image reconstruction method is urgently needed. Summary of the Invention
[0003] The purpose of the present invention is to overcome one or more of the above-mentioned existing technical problems and provide a dual-branch deep compressed sensing reconstruction method based on CNN and Mamba.
[0004] To achieve the above objectives, the present invention provides a dual-branch deep compressed sensing reconstruction method based on CNN and Mamba, comprising: Get a high-dimensional input image; Linearly project the high-dimensional input image through the sampling matrix to obtain low-dimensional measurements; The low-dimensional measurements are fed into the initial reconstruction module to obtain a high-resolution preliminary feature map. Input the low-dimensional measurement value into the initial feature extraction module to obtain a high-resolution initial feature map; Input the high-resolution initial feature map into the CNN branch feature extraction module to obtain a local feature set; The local feature set and the high-resolution initial feature map are input into the state-aware transformation module to obtain the global context-related output features; Recombining the global context-related output features and the high-resolution preliminary feature map to obtain recombined context features and a recombined feature map; The reconstructed context features, reconstructed feature maps and local feature sets are input into the final image reconstruction module to obtain a high-quality reconstructed image.
[0005] According to one aspect of the present invention, the method for obtaining a high-resolution preliminary feature map is: Input the low-dimensional measurements into the initial reconstruction module; The low-dimensional measurement values are reverse-sampled based on the transposed sampling matrix to obtain a preliminary high-dimensional estimate, where the formula is, ; in, represents a preliminary high-dimensional estimate; represents the transpose of the measurement matrix; The preliminary high-dimensional estimate is upsampled based on the pixel rearrangement function to obtain a high-resolution preliminary feature map, where the formula is, ; in, Represents a high-resolution preliminary feature map; Represents the pixel rearrangement function.
[0006] According to one aspect of the present invention, the method for obtaining a high-resolution initial feature map is: The low-dimensional measurement value is input into the initial feature extraction module, and the high-resolution initial feature map is extracted from the low-dimensional measurement value through convolution and upsampling operations. The formula is: ; in, Represents the high-resolution initial feature map; Represents PixelShuffle upsampling operation; Represents a 1×1 convolution operation.
[0007] According to one aspect of the present invention, the method for obtaining the local feature set is: Input the high-resolution initial feature map into the CNN branch feature extraction module; After the high-resolution initial feature map and the high-resolution initial feature map operated by the i-th convolution block are simultaneously input into the i-th residual block for feature extraction, they are further upsampled to obtain a local feature set containing four different scale features, where each feature is a specific representation of the local information extracted at different resolutions, and the detailed structure of the high-resolution initial feature map is retained. The formula is, ; ; ; in, represents the i-th scale feature; represents the upsampling operation of the i-th layer; Indicates feature extraction through the i-th residual block; Represents the i-th layer convolution block operation; Represents a set of local features; represents the first scale feature; Represents the second scale feature; Represents the third scale feature; Represents the fourth scale feature.
[0008] According to one aspect of the present invention, the method for obtaining global context-related output features is: The state-aware transformation module includes four state-space models and an attention mechanism unified module; The high-resolution initial feature map is input into the unified module of the first state space model and the attention mechanism to obtain the output feature of the unified module of the first state space model and the attention mechanism, where the formula is, ; in, Represents the output features of the first state space model and the unified module of the attention mechanism; Represents the first state space model and attention mechanism unified module; The output features of the first state space model and attention mechanism unified module are input into the second state space model and attention mechanism unified module and combined with the first scale features to obtain the output features of the second state space model and attention mechanism unified module, where the formula is, ; in, Represents the unified module of the second state space model and the attention mechanism; Represents the output features of the unified module of the second state space model and the attention mechanism; The output features of the second state space model and attention mechanism unified module are input into the third state space model and attention mechanism unified module and combined with the second scale features to obtain the output features of the third state space model and attention mechanism unified module, where the formula is, ; in, Represents the unified module of the third state space model and the attention mechanism; Represents the output features of the unified module of the third state space model and the attention mechanism; The output features of the second state space model and attention mechanism unified module are input into the fourth state space model and attention mechanism unified module and combined with the third scale features and the fourth scale features to obtain the global context-related output features, where the formula is, ; in, Represents the fourth state space model and attention mechanism unified module; Represents the global context-dependent output features.
[0009] According to one aspect of the present invention, the method for obtaining the recombinant context features and the recombinant feature graph is: Based on the reorganization function, a unified reorganization operation is performed on the global context-related output features and the high-resolution preliminary feature map, and the feature tensor in the sequence format is reorganized into a two-dimensional representation that conforms to the image space format, so that the features are aligned and fused in the spatial domain, and the reorganized context features and reorganized feature maps are obtained. The formula is, ; ; in, Represents reorganization context features; Represents the recombination feature map; represents the reorganization function; Indicates the batch size; Indicates the height of the image; Indicates the width of the image.
[0010] According to one aspect of the present invention, the method for obtaining a high-quality reconstructed image is: The reconstructed context features, reconstructed feature maps and local feature sets are fused in the spatial dimension, and comprehensive information is extracted through deep convolution and transformed into a high-quality reconstructed image, where the formula is, ; in, Indicates high-quality reconstructed image; Represents the fusion extraction module.
[0011] According to one aspect of the present invention, the quality of the generated high-quality reconstructed image is optimized based on a loss function, and the coordinated expression of global context and local detail features is obtained by minimizing the error between the generated image and the real image, wherein the formula of the loss function is: ; in, represents the loss function.
[0012] To achieve the above objectives, the present invention provides a dual-branch deep compressed sensing reconstruction system based on CNN and Mamba, comprising: High-dimensional input image acquisition module: obtains high-dimensional input images; Low-dimensional measurement value generation module: linearly project the high-dimensional input image through the sampling matrix to obtain low-dimensional measurement values; High-resolution preliminary feature map generation module: low-dimensional measurements are input into the initial reconstruction module to obtain a high-resolution preliminary feature map; High-resolution initial feature map generation module: Input the low-dimensional measurement value into the initial feature extraction module to obtain a high-resolution initial feature map; Local feature set generation module: Input the high-resolution initial feature map into the CNN branch feature extraction module to obtain the local feature set; Global context-related output feature generation module: The local feature set and the high-resolution initial feature map are input into the state-aware transformation module to obtain the global context-related output features; Recombined feature generation module: recombines the global context-related output features and the high-resolution preliminary feature map to obtain recombined context features and recombined feature maps; High-quality reconstructed image generation module: The reconstructed context features, reconstructed feature maps and local feature sets are input into the final image reconstruction module to obtain a high-quality reconstructed image.
[0013] Based on this, the beneficial effects of the present invention are as follows: This application proposes a dual-branch architecture combining a state-space model and a CNN for compressed sensing image reconstruction, which can adapt to the needs of different scenarios. It designs a sequence modeling branch SATM based on Mamba, uses its efficient selective state-space mechanism to capture the global dependency features of the image, and constructs a complementary CNN feature extraction branch, which focuses on extracting local spatial detail information and realizes the effective fusion of global and local features. This application designs a multi-scale feature fusion mechanism that implements feature interaction and integration at different levels. It fully utilizes the SATM branch's ability to model long-range dependencies and the CNN branch's ability to extract local features. It can maintain high reconstruction quality, effectively integrate global and local features, and improve the reconstruction effect of complex image structures. The model effectively completes the complete conversion process from local to global, from low-dimensional measurements to high-resolution reconstructed images. The final image not only contains the key information of the original measurements, but also combines global context and multi-scale local features, thereby significantly improving the reconstruction effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 This is a flowchart of a dual-branch deep compressed sensing reconstruction method based on CNN and Mamba according to an exemplary embodiment; Figure 2 1 is a diagram showing the overall network structure of a dual-branch deep compressed sensing reconstruction method based on CNN and Mamba according to an exemplary embodiment; Figure 3 is a comparison diagram of a dual-branch deep compressed sensing reconstruction method based on CNN and Mamba and other methods according to an exemplary embodiment; Figure 4 The present invention is a flowchart of a dual-branch deep compressed sensing reconstruction system based on CNN and Mamba according to an exemplary embodiment. DETAILED DESCRIPTION
[0015] The present invention will now be discussed with reference to exemplary embodiments. It should be understood that the embodiments discussed are only for enabling those skilled in the art to better understand and thereby implement the present invention, rather than implying any limitation on the scope of the present invention.
[0016] As used herein, the term “including” and variations thereof are to be interpreted as open-ended terms meaning “including, but not limited to.” The term “based on” is to be interpreted as “based, at least in part, on,” and the terms “one embodiment” and “an embodiment” are to be interpreted as “at least one embodiment.”
[0017] According to one embodiment of the present invention, Figure 1 FIG. 1 is a flowchart of a dual-branch deep compressed sensing reconstruction method based on CNN and Mamba according to an exemplary embodiment. Figure 1 As shown, to achieve the above purpose, the present invention provides a dual-branch deep compressed sensing reconstruction method based on CNN and Mamba, comprising: Get a high-dimensional input image; The high-dimensional input image is linearly projected through the sampling matrix to obtain a low-dimensional measurement value. The dimension of the measurement value is much smaller than the dimension of the input signal. The main purpose of this stage is to compress the high-dimensional input image to a low-dimensional measurement value to reduce the computational burden of subsequent processing; The low-dimensional measurements are fed into the initial reconstruction module to obtain a high-resolution preliminary feature map. Input the low-dimensional measurement value into the initial feature extraction module to obtain a high-resolution initial feature map; Input the high-resolution initial feature map into the CNN branch feature extraction module to obtain a local feature set; The local feature set and the high-resolution initial feature map are input into the state-aware transformation module to obtain the global context-related output features; Recombining the global context-related output features and the high-resolution preliminary feature map to obtain recombined context features and a recombined feature map; The reconstructed context features, reconstructed feature maps and local feature sets are input into the final image reconstruction module to obtain a high-quality reconstructed image.
[0018] Figure 2 FIG is an overall network structure diagram of a dual-branch deep compressed sensing reconstruction method based on CNN and Mamba according to an exemplary embodiment. Figure 2 As shown, according to one embodiment of the present invention, the method for obtaining a high-resolution preliminary feature map is: In the initial reconstruction stage, the high-dimensional preliminary feature map is restored from the low-dimensional measurement value. This stage is divided into two steps: desampling operation and upsampling operation. In order to convert the measurement value back to high-dimensional features, the transposed sampling matrix is used for desampling to obtain a preliminary high-dimensional estimate: Input the low-dimensional measurements into the initial reconstruction module; The low-dimensional measurement values are reverse-sampled based on the transposed sampling matrix to obtain a preliminary high-dimensional estimate, where the formula is, ; in, represents a preliminary high-dimensional estimate; represents the transpose of the measurement matrix; The preliminary high-dimensional estimate is upsampled based on the pixel rearrangement function to obtain a high-resolution preliminary feature map, where the formula is, ; in, Represents a high-resolution preliminary feature map; Represents the pixel rearrangement function.
[0019] Through these two steps, a high-resolution preliminary feature map is generated from the low-dimensional measurements, which provides input features for the subsequent feature extraction and fusion processes.
[0020] According to one embodiment of the present invention, a method for obtaining a high-resolution initial feature map is: The low-dimensional measurement value is input into the initial feature extraction module, and the high-resolution initial feature map is extracted from the low-dimensional measurement value through convolution and upsampling operations. The formula is: ; in, Represents the high-resolution initial feature map; Represents PixelShuffle upsampling operation; Represents a 1×1 convolution operation.
[0021] A high-resolution initial feature map is generated from the measurements, providing high-quality input for the subsequent processing of CNN64 and SATM branches.
[0022] According to one embodiment of the present invention, the method for obtaining the local feature set is: In order to further extract local details and multi-scale information from the initial features, the CNN branch generator module is designed to extract and expand features layer by layer, thereby generating a set of local features suitable for subsequent global modeling. Input the high-resolution initial feature map into the CNN branch feature extraction module; The high-resolution initial feature map and the high-resolution initial feature map operated by the i-th convolution block are simultaneously input to the i-th residual block for feature extraction to retain the input information and enhance the details. The upsampling operation is continued to obtain a local feature set containing four different scale features. The input features are transformed by stacking convolution layers to extract rich local information. After multi-level calculations of the CNN64 module, feature maps of different scales (such as 8x8, 16x16, 32x32, 64x64) are gradually generated, where each feature is a specific representation of the local information extracted at different resolutions, and the detailed structure of the high-resolution initial feature map is retained. The formula is, ; ; ; in, represents the i-th scale feature; represents the upsampling operation of the i-th layer; Indicates feature extraction through the i-th residual block; Represents the i-th layer convolution block operation; Represents a set of local features; represents the first scale feature; Represents the second scale feature; Represents the third scale feature; Represents the fourth scale feature.
[0023] According to one embodiment of the present invention, the method for obtaining the output features related to the global context is: To further capture global context and long-range dependencies, we designed the SATM module, which combines the traditional Transformer architecture with the innovative state-space model Mamba2. The state-aware transformation module includes four state-space models and an attention mechanism unified module. The high-resolution initial feature map is input into the unified module of the first state space model and the attention mechanism to obtain the output feature of the unified module of the first state space model and the attention mechanism, where the formula is, ; in, Represents the output features of the first state space model and the unified module of the attention mechanism; Represents the first state space model and attention mechanism unified module; The output features of the first state space model and attention mechanism unified module are input into the second state space model and attention mechanism unified module and combined with the first scale features to obtain the output features of the second state space model and attention mechanism unified module, where the formula is, ; in, Represents the unified module of the second state space model and the attention mechanism; Represents the output features of the unified module of the second state space model and the attention mechanism; The output features of the second state space model and attention mechanism unified module are input into the third state space model and attention mechanism unified module and combined with the second scale features to obtain the output features of the third state space model and attention mechanism unified module, where the formula is, ; in, Represents the unified module of the third state space model and the attention mechanism; Represents the output features of the unified module of the third state space model and the attention mechanism; The output features of the second state space model and attention mechanism unified module are input into the fourth state space model and attention mechanism unified module and combined with the third scale features and the fourth scale features to obtain the global context-related output features, where the formula is, ; in, Represents the fourth state space model and attention mechanism unified module; Represents the global context-dependent output features.
[0024] According to one embodiment of the present invention, the method for obtaining the recombinant context feature and the recombinant feature graph is: The SATM module generates output features through global modeling. These features combine multi-scale local features with long-range global context information. At the same time, the preliminary features retain important information in the original measurements. In the feature fusion stage, these two types of features need to be uniformly recombined to form a representation with consistent spatial dimensions, providing a complete input for the subsequent image reconstruction process. Based on the reorganization function, a unified reorganization operation is performed on the global context-related output features and the high-resolution preliminary feature map, and the feature tensor in the sequence format is reorganized into a two-dimensional representation that conforms to the image space format, so that the features are aligned and fused in the spatial domain, and the reorganized context features and reorganized feature maps are obtained. The formula is, ; ; in, Represents reorganization context features; Represents the recombination feature map; represents the reorganization function; Indicates the batch size; Indicates the height of the image; Indicates the width of the image.
[0025] Through the above feature recombination operation, the features and are converted into a consistent high-dimensional spatial representation. This stage not only integrates global and local information but also provides complete and high-quality input features for final image generation, bridging the key link between feature modeling and reconstruction tasks.
[0026] According to one embodiment of the present invention, a method for obtaining a high-quality reconstructed image is: The reconstructed context features, reconstructed feature maps and local feature sets are fused in the spatial dimension, and comprehensive information is extracted through deep convolution and transformed into a high-quality reconstructed image, where the formula is, ; in, Indicates high-quality reconstructed image; Represents the fusion extraction module.
[0027] Through the above operations, the model effectively completes the complete conversion process from local to global, from low-dimensional measurements to high-resolution reconstructed images. The final image not only contains the key information of the original measurement values, but also combines global context and multi-scale local features, thereby significantly improving the reconstruction effect and providing an efficient and innovative solution for high-precision image restoration in the field of compressed sensing.
[0028] According to one embodiment of the present invention, the quality of the generated high-quality reconstructed image is optimized based on a loss function, and the coordinated expression of global context and local detail features is obtained by minimizing the error between the generated image and the real image, wherein the formula of the loss function is: ; in, represents the loss function.
[0029] Through the above loss function, the model effectively combines the key information of global features, local features and original measurements in the feature fusion stage, thereby ensuring that the final generated image not only contains clear detailed textures but also has the accuracy of global structure. This design further enhances the model's adaptability to complex image scenes and lays the foundation for high-quality output of the final image reconstruction.
[0030] Figure 3 is a comparison diagram of a dual-branch deep compressed sensing reconstruction method based on CNN and Mamba and other methods according to an exemplary embodiment. Figure 3 As shown, according to one embodiment of the present invention, the training data of the present method is based on the BSD500 dataset
[18] , which contains three image sets: 200 for training, 100 for validation, and 200 for testing. Set11 is used as the validation dataset. 200 96×96 pixel image patches are extracted from each training image, resulting in a total of 100,000 training samples. Image diversity is increased through data augmentation techniques, including bidirectional flipping, rotation operations, and scale adjustment. Performance evaluation is performed on the BSD100
[19] , Set5
[20] , Urban100
[21] , and UCMerced benchmark datasets. Among them, BSD500, Set5, and Urban100 are widely used in the field of CS research and provide standardized and diverse benchmarks for evaluating reconstruction methods. These datasets contain rich structural patterns, various textures, and complex details, which are crucial for evaluating the generalization ability of the model on different data types. In addition, the introduction of the UCMerced dataset expands the test scope, allowing the evaluation to cover remote sensing image reconstruction tasks, further verifying the adaptability of the present method in practical application scenarios. To ensure fairness, all methods are trained and tested on the same dataset when compared with state-of-the-art methods (ISTA-Net+, CSNet, and AMP-Net). This standardization avoids potential bias and provides a reliable basis for performance benchmarking.
[0031] The proposed model is compared with several leading methods, such as ISTA-Net+, CsNet, and AMP-Net, which combine traditional algorithms with deep learning techniques. The evaluation is based on perceptual metrics, namely peak signal-to-noise ratio (PSNR)
[22] and structural similarity (SSIM)
[23] , where PSNR is used to measure the quality of the reconstructed image and SSIM is used to evaluate structural similarity. The higher the value of these two metrics, the better the performance. These models were obtained from their respective sources and run with the default configuration. To ensure fairness, the training images of all competing models were selected from the BSD500 dataset. The training parameters of this method are set as follows: the total training rounds are 50 epochs, the initial learning rate is set to 0.0001, and the optimizer uses the Adam optimizer with parameters The experiments are conducted on a machine equipped with an Intel Xeon 8336 CPU and a GeForce RTX 4090 GPU.
[0032] For comparison purposes, our method is comprehensively and quantitatively compared with current mainstream methods (ISTA-Net+, CSNet, and AMP-Net) on four benchmark datasets (Set5, Urban100, BSD100, and UCMerecd) at different sampling rates (0.04-0.50). On the Set5 dataset, our method consistently outperforms other methods at all sampling rates. In particular, at τ = 0.50, our method achieves a PSNR of 42.02 dB and an SSIM of 0.9811, significantly outperforming the next-best AMP-Net (41.48 dB / 0.9756). This performance advantage is even more pronounced at a lower sampling rate (τ = 0.04), where our method achieves a PSNR of 28.02 dB, at least 0.77 dB higher than other methods. On the BSD100 dataset, our method demonstrates comprehensive performance advantages. At τ=0.50, a PSNR of 36.60dB and an SSIM of 0.9632 were achieved, surpassing all compared methods. In particular, at low sampling rates (τ=0.04), our method (25.94dB) improved by nearly 2dB compared to the second-best AMP-Net (24.04dB), demonstrating strong reconstruction capabilities in challenging low sampling rate scenarios.
[0033] On the Urban100 dataset, this method demonstrates excellent PSNR performance, particularly at low sampling rates. For example, at a sampling rate of τ = 0.5, this method achieves a PSNR of 34.77 dB, exceeding other methods and demonstrating the advantages of pixel-level reconstruction. A high PSNR value accurately restores image details, visually preserving sharp edges and subtle features of geometric shapes, making the reconstructed image closer to the original image in terms of detail.
[0034] This method also demonstrates excellent reconstruction performance on the UCMerced remote sensing dataset. Even at a low sampling rate (τ = 0.04), it maintains a significant advantage, achieving a PSNR of 25.23dB and an SSIM of 0.6827, an improvement of approximately 1.36dB over the next-best method, AMP-Net (23.87dB / 0.6543). Overall, this method's excellent performance on the UCMerced dataset demonstrates its broad potential for application beyond natural image reconstruction tasks in remote sensing image reconstruction.
[0035] This method adopts a dual-branch architecture, with the SATM branch modeling global contextual dependencies and the CNN branch extracting local features. This design efficiently fuses multi-scale features at low sampling rates, significantly improving PSNR. However, the emphasis on fine-grained recovery of pixel-level details may weaken global structural consistency, resulting in a slightly lower SSIM. Although the SSIM values for some sampling rates are suboptimal, this method still achieves suboptimal results, demonstrating its excellent performance in terms of global perceptual consistency.
[0036] Especially in complex textured images, this method achieves a PSNR of 22.63dB at τ = 0.04, nearly 2dB higher than AMP-Net, and its ability to restore detail is particularly outstanding. Despite the lower SSIM, the advantage of PSNR lies in more accurately restoring local visual details, allowing images to still present good visual effects at low sampling rates.
[0037] These experimental results clearly demonstrate that our method achieves optimal or near-optimal performance across a wide range of datasets and sampling rates. In particular, our method demonstrates significant advantages in challenging low-sampling scenarios, confirming its effectiveness in image reconstruction tasks.
[0038] In addition, to achieve the above-mentioned purpose, the present invention also provides a dual-branch deep compressed sensing reconstruction system based on CNN and Mamba. Figure 4 FIG. 1 is a flowchart of a dual-branch deep compressed sensing reconstruction system based on CNN and Mamba according to an exemplary embodiment. Figure 4 As shown, a dual-branch deep compressed sensing reconstruction system based on CNN and Mamba in the present invention includes: High-dimensional input image acquisition module: obtains high-dimensional input images; Low-dimensional measurement value generation module: linearly project the high-dimensional input image through the sampling matrix to obtain low-dimensional measurement values; High-resolution preliminary feature map generation module: low-dimensional measurements are input into the initial reconstruction module to obtain a high-resolution preliminary feature map; High-resolution initial feature map generation module: Input the low-dimensional measurement value into the initial feature extraction module to obtain a high-resolution initial feature map; Local feature set generation module: Input the high-resolution initial feature map into the CNN branch feature extraction module to obtain the local feature set; Global context-related output feature generation module: The local feature set and the high-resolution initial feature map are input into the state-aware transformation module to obtain the global context-related output features; Recombined feature generation module: recombines the global context-related output features and the high-resolution preliminary feature map to obtain recombined context features and recombined feature maps; High-quality reconstructed image generation module: The reconstructed context features, reconstructed feature maps and local feature sets are input into the final image reconstruction module to obtain a high-quality reconstructed image.
[0039] Those skilled in the art will appreciate that the modules and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented using electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0040] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and equipment can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0041] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0042] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of these modules may be selected according to actual needs to achieve the objectives of the embodiments of the present invention.
[0043] In addition, each functional module in the embodiment of the present invention may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0044] If the functions are implemented as software modules and sold or used as standalone products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or the portion of the technical solution itself, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the energy-saving signal transmission / reception method according to various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, ROM, RAM, a magnetic disk, or an optical disk.
[0045] The above description is merely a preferred embodiment of the present application and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the invention herein is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the inventive concept. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features having similar functions disclosed in this application.
[0046] It should be understood that the size of the serial numbers of each step in the content of the invention and the embodiments of the present invention does not absolutely mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
Claims
1. A dual-branch deep compressed sensing reconstruction method based on CNN and Mamba, characterized by: include: Get a high-dimensional input image; Linearly project the high-dimensional input image through the sampling matrix to obtain low-dimensional measurements; The low-dimensional measurements are fed into the initial reconstruction module to obtain a high-resolution preliminary feature map. Input the low-dimensional measurement value into the initial feature extraction module to obtain a high-resolution initial feature map; Input the high-resolution initial feature map into the CNN branch feature extraction module to obtain a local feature set; The local feature set and the high-resolution initial feature map are input into the state-aware transformation module to obtain the global context-related output features; Recombining the global context-related output features and the high-resolution preliminary feature map to obtain recombined context features and a recombined feature map; The reconstructed context features, reconstructed feature maps and local feature sets are input into the final image reconstruction module to obtain a high-quality reconstructed image.
2. The dual-branch deep compressed sensing reconstruction method based on CNN and Mamba according to claim 1, characterized in that: The method for obtaining a high-resolution preliminary feature map is: Input the low-dimensional measurements into the initial reconstruction module; The low-dimensional measurement values are reverse-sampled based on the transposed sampling matrix to obtain a preliminary high-dimensional estimate, where the formula is, ; in, represents a preliminary high-dimensional estimate; represents the transpose of the measurement matrix; The preliminary high-dimensional estimate is upsampled based on the pixel rearrangement function to obtain a high-resolution preliminary feature map, where the formula is, ; in, Represents a high-resolution preliminary feature map; Represents the pixel rearrangement function.
3. The dual-branch deep compressed sensing reconstruction method based on CNN and Mamba according to claim 2, characterized in that: The method for obtaining a high-resolution initial feature map, The low-dimensional measurement value is input into the initial feature extraction module, and the high-resolution initial feature map is extracted from the low-dimensional measurement value through convolution and upsampling operations. The formula is: ; in, Represents the high-resolution initial feature map; Represents PixelShuffle upsampling operation; Represents a 1×1 convolution operation.
4. The dual-branch deep compressed sensing reconstruction method based on CNN and Mamba according to claim 3, characterized in that: The method for obtaining the local feature set is: Input the high-resolution initial feature map into the CNN branch feature extraction module; After the high-resolution initial feature map and the high-resolution initial feature map operated by the i-th convolution block are simultaneously input into the i-th residual block for feature extraction, they are further upsampled to obtain a local feature set containing four different scale features, where each feature is a specific representation of the local information extracted at different resolutions, and the detailed structure of the high-resolution initial feature map is retained. The formula is, ; ; ; in, represents the i-th scale feature; represents the upsampling operation of the i-th layer; Indicates feature extraction through the i-th residual block; Represents the i-th layer convolution block operation; Represents a set of local features; represents the first scale feature; Represents the second scale feature; Represents the third scale feature; Represents the fourth scale feature.
5. The dual-branch deep compressed sensing reconstruction method based on CNN and Mamba according to claim 4, characterized in that: The method for obtaining the output features related to the global context is: The state-aware transformation module includes four state-space models and an attention mechanism unified module; The high-resolution initial feature map is input into the unified module of the first state space model and the attention mechanism to obtain the output feature of the unified module of the first state space model and the attention mechanism, where the formula is, ; in, Represents the output features of the first state space model and the unified module of the attention mechanism; Represents the first state space model and attention mechanism unified module; The output features of the first state space model and attention mechanism unified module are input into the second state space model and attention mechanism unified module and combined with the first scale features to obtain the output features of the second state space model and attention mechanism unified module, where the formula is, ; in, Represents the unified module of the second state space model and the attention mechanism; Represents the output features of the unified module of the second state space model and the attention mechanism; The output features of the second state space model and attention mechanism unified module are input into the third state space model and attention mechanism unified module and combined with the second scale features to obtain the output features of the third state space model and attention mechanism unified module, where the formula is, ; in, Represents the unified module of the third state space model and the attention mechanism; Represents the output features of the unified module of the third state space model and the attention mechanism; The output features of the second state space model and attention mechanism unified module are input into the fourth state space model and attention mechanism unified module and combined with the third scale features and the fourth scale features to obtain the global context-related output features, where the formula is, ; in, Represents the fourth state space model and attention mechanism unified module; Represents the global context-dependent output features.
6. The dual-branch deep compressed sensing reconstruction method based on CNN and Mamba according to claim 5, characterized in that: The method for obtaining the recombinant context features and the recombinant feature graph is: Based on the reorganization function, a unified reorganization operation is performed on the global context-related output features and the high-resolution preliminary feature map, and the feature tensor in the sequence format is reorganized into a two-dimensional representation that conforms to the image space format, so that the features are aligned and fused in the spatial domain, and the reorganized context features and reorganized feature maps are obtained. The formula is, ; ; in, Represents reorganization context features; Represents the recombination feature map; represents the reorganization function; Indicates the batch size; Indicates the height of the image; Indicates the width of the image.
7. The dual-branch deep compressed sensing reconstruction method based on CNN and Mamba according to claim 6, characterized in that: The method for obtaining a high-quality reconstructed image is: The reconstructed context features, reconstructed feature maps and local feature sets are fused in the spatial dimension, and comprehensive information is extracted through deep convolution and transformed into a high-quality reconstructed image, where the formula is, ; in, Indicates high-quality reconstructed image; Represents the fusion extraction module.
8. The dual-branch deep compressed sensing reconstruction method based on CNN and Mamba according to claim 7, characterized in that: It also includes optimizing the quality of generating high-quality reconstructed images based on the loss function, and obtaining the coordinated expression of global context and local detail features by minimizing the error between the generated image and the real image. The formula of the loss function is, ; in, represents the loss function.
9. A dual-branch deep compressed sensing reconstruction system based on CNN and Mamba, characterized by: include: High-dimensional input image acquisition module: obtains high-dimensional input images; Low-dimensional measurement value generation module: linearly project the high-dimensional input image through the sampling matrix to obtain low-dimensional measurement values; High-resolution preliminary feature map generation module: low-dimensional measurements are input into the initial reconstruction module to obtain a high-resolution preliminary feature map; High-resolution initial feature map generation module: Input the low-dimensional measurement value into the initial feature extraction module to obtain a high-resolution initial feature map; Local feature set generation module: Input the high-resolution initial feature map into the CNN branch feature extraction module to obtain the local feature set; Global context-related output feature generation module: The local feature set and the high-resolution initial feature map are input into the state-aware transformation module to obtain the global context-related output features; Recombined feature generation module: recombines the global context-related output features and the high-resolution preliminary feature map to obtain recombined context features and recombined feature maps; High-quality reconstructed image generation module: The reconstructed context features, reconstructed feature maps and local feature sets are input into the final image reconstruction module to obtain a high-quality reconstructed image.
Citation Information
Cited By
Monitoring image super-division reconstruction method based on frequency domain decoupling and rotation perception Mangbar
CN121903842A