Image super-resolution reconstruction method and system under blurred field of view in deep earth disaster environment
By combining the progressive Mamba network with the progressive convolution chain and the state-space module, the robustness and efficiency issues of image super-resolution reconstruction in deep-earth disaster environments are solved, and efficient image reconstruction effects are achieved, which is suitable for edge devices.
Patent Information
- Application Number
- CN202411892254.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2044-12-20
AI Technical Summary
Existing image super-resolution reconstruction technology has poor robustness in deep-earth disaster environments, and it is difficult to balance the number of parameters, computational complexity, and reconstruction performance. In addition, existing methods are difficult to deploy efficiently on edge devices.
The progressive Mamba network is adopted, combined with progressive convolution chains and state-space modules, and shallow and deep feature extraction modules are designed. Image reconstruction is achieved through channel shuffling and sub-pixel upsampling.
It achieves a balance between the number of model parameters, computational complexity, and reconstruction performance, improves the efficiency and quality of image reconstruction, and is suitable for edge device deployment.
Smart Images

Figure CN120198287B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image super-resolution reconstruction, and in particular relates to a method and system for image super-resolution reconstruction in a fuzzy field of view in a deep-earth disaster environment. Background Art
[0002] In deep disaster environments, poor lighting, dusty, or smoky conditions significantly degrade surveillance image quality, blurring image details and causing significant noise interference, compromising their practical application. These issues not only pose challenges to post-disaster monitoring, rescue command, and personnel location, but also significantly limit the effectiveness of modern disaster recovery systems. Therefore, efficiently restoring low-quality images in disaster environments has become a crucial prerequisite for improving emergency response capabilities and ensuring the success of rescue missions.
[0003] Image super-resolution reconstruction technology uses "soft processing" to reconstruct low-resolution images into high-resolution ones, compensating for the loss of image detail. This technology not only provides clear input images for applications such as object detection and path planning, but also effectively improves the performance of the entire monitoring system. In post-disaster recovery and emergency response, super-resolution technology can significantly enhance image quality, thereby improving the accuracy and efficiency of rescue decisions.
[0004] Image super-resolution technology can be roughly divided into three categories: interpolation-based methods, reconstruction-based methods, and learning-based methods. Interpolation-based methods are one of the earliest forms of image super-resolution technology. This method estimates the new pixel values in the high-resolution image by interpolating the pixels in the low-resolution image. Common interpolation techniques include nearest neighbor interpolation, bilinear interpolation, and bicubic interpolation. The advantages of these methods are simple calculations, fast speed, and suitability for real-time applications. However, their main disadvantage is that the reconstructed images usually lack details and have blurred edges, making it difficult to restore the delicate features of the real image. In particular, when dealing with complex patterns and textures, the interpolation method often has unsatisfactory results.
[0005] Reconstruction-based methods restore high-resolution images by modeling prior knowledge of the image. These methods typically exploit the coherence of multiple frames or achieve super-resolution reconstruction through sparse representation of the image. Common reconstruction techniques include transform-domain-based sparse representation methods and statistical learning methods. These methods are relatively good at restoring details and improving image quality, and can better preserve image edge and texture information. However, they have high computational complexity and long processing time, making them less efficient, especially when processing large amounts of data.
[0006] In recent years, learning-based methods, particularly deep learning methods, have become a hot topic in image super-resolution research. These methods typically utilize deep learning models such as convolutional neural networks to learn the mapping relationship from low-resolution images to high-resolution images through training on large-scale image datasets. Learning-based methods have strong learning capabilities and can automatically extract complex features in images, thereby generating high-quality high-resolution images. These methods have performed well in multiple image super-resolution challenges, effectively restoring details and improving the visual quality of reconstructed images.
[0007] Although image super-resolution reconstruction technology has achieved certain applications, existing methods still suffer from several significant drawbacks. In practical applications, images may be affected by various factors such as noise, blur, and compression. Existing super-resolution reconstruction methods are poorly robust when dealing with these interferences, resulting in suboptimal reconstruction results. While existing image super-resolution reconstruction networks have significantly improved reconstruction performance, they struggle to balance the number of parameters, computational complexity, and reconstruction performance. Training is challenging and deployment on edge devices is difficult. Summary of the Invention
[0008] In order to solve the above technical problems, the present invention provides a method and system for super-resolution reconstruction of images under blurred fields of view in deep-earth disaster environments. The method proposes a progressive Mamba network. The design of this network fully considers the important role of large receptive fields in image reconstruction tasks and can effectively capture global information in images. Through this structure, a good balance between the number of model parameters, the amount of calculation and the reconstruction performance is successfully achieved, so that the network can be more efficient in practical applications while maintaining high performance. In addition, the present invention attaches importance to the guiding role of shallow features in image reconstruction and designs a shallow feature extraction module organized by a dual-path convolution chain, thereby improving the adaptability of the network model to various image features.
[0009] In order to achieve the above object, the present invention is achieved through the following technical solutions:
[0010] The present invention provides a method for super-resolution reconstruction of images in a blurred field of view in a deep earth disaster environment. The method is characterized in that: the method is implemented by a progressive Mamba network, which includes a shallow feature extraction module composed of two progressive convolution chains with opposite convolution orders, a deep feature extraction module, and a reconstruction module composed of a channel shuffling strategy and a sub-pixel upsampling module. The specific reconstruction method includes the following steps:
[0011] Step 1: Take the high-resolution image I HR Perform bicubic downsampling to obtain a low-resolution image I LR , the low-resolution image I is obtained by the shallow feature extraction moduleLR Perform shallow feature extraction. The extraction process is described as follows:
[0012] F S =H Shallow (I LR )
[0013] Among them, H Shallow (·) is the shallow feature extraction module, F S Output features of the shallow feature extraction module;
[0014] Step 2: Input the output features of the shallow feature extraction module into the deep feature extraction module to extract deep features, wherein the deep feature extraction module includes four branches of three progressive state space modules PRSSB, each branch processes the output features of the input shallow feature extraction module and outputs an output feature respectively;
[0015] Step 3: Fuse the output features of the four branches in step 2 together, and use the reconstruction module to reconstruct the fused features to obtain the final high-resolution reconstructed image I SR .
[0016] A further improvement of the present invention is that the two progressive convolution chains with opposite convolution orders in the shallow feature extraction module are an upper branch convolution chain and a lower branch convolution chain, and the upper branch convolution chain and the lower branch convolution chain respectively include a convolution layer with a convolution kernel size of 5, a convolution layer with a convolution kernel size of 3, and a convolution layer with a convolution kernel size of 1.
[0017] A further improvement of the present invention is that: in step 1, the upper branch convolution chain and the lower branch convolution chain are used to perform the low-resolution image I LR The processing process is:
[0018] F S1 =C1(C3(C5(I LR )));
[0019] F S2 =C5(C3(C1(I LR )));
[0020] F S =C1(CShuffle(Concat(F S1 ,F S2 )))
[0021] Among them, C5 represents the convolution layer with a convolution kernel size of 5, C3 represents the convolution layer with a convolution kernel size of 3, C1 represents the convolution layer with a convolution kernel size of 1, and F S1 is the output result of the upper branch convolution chain, F S2is the output result of the convolution chain of the lower branch, Concat() represents splicing by channel dimension, CShuffle() represents the channel shuffling strategy, F S It is the output feature of the shallow feature extraction module.
[0022] A further improvement of the present invention is that the first branch in the deep-level feature extraction module includes a convolution layer with a convolution kernel size of 5, a convolution layer with a convolution kernel size of 3, and a convolution layer with a convolution kernel size of 1. In the first branch, a convolution chain that gradually adjusts the convolution kernel size is adopted to achieve accurate extraction of image features. At the same time, through multiple upsampling operations, the first branch 1 can effectively increase the image size, thereby retaining more detail information and contextual features. The second branch in the deep-level feature extraction module includes three progressive state space modules PRSSB, the third branch in the deep-level feature extraction module includes a convolution layer with a convolution kernel size of 5, an upsampling operation, and a convolution layer with a convolution kernel size of 3. The fourth branch in the deep-level feature extraction module is a convolution layer with a convolution kernel size of 5.
[0023] A further improvement of the present invention is that the output features of the input shallow feature extraction module are processed by the deep feature extraction module, and the specific method of extracting deep features is:
[0024] Step 2.1: Extract the output feature F of the shallow feature extraction module from the input through the first branch S Processing is performed to obtain the output feature F of the first branch D1 :
[0025] F D1 =C1(UP(C3(UP(C5(F S )))))
[0026] Step 2.2: The output feature F of the shallow feature extraction module of the input is extracted through the second branch S After upsampling and processing by the progressive state space module PRSSB, the output feature F of the second branch is obtained D2 :
[0027] F D2 =PRSSB(UP(PRSSB(UP(PRSSB(F S )))))
[0028] Step 2.3: Upsample the output of the first progressive state space module PRSSB in the second branch through the third molecule to obtain the output feature F of the third branch. D3 :
[0029] F D3=C3(UP(C5(UP(PRSSB(F S )))))
[0030] Step 2.4: After upsampling the output result of the second progressive state space module PRSSB in the second branch through the fourth branch, the output feature F of the fourth branch is obtained. D4 :
[0031] F D4 =C5(UP(PRSSB(UP(PRSSB(F S )))))
[0032] Among them, F S is the output feature of the shallow feature extraction module, C5 represents the convolution layer with a convolution kernel size of 5, C3 represents the convolution layer with a convolution kernel size of 3, C1 represents the convolution layer with a convolution kernel size of 1, UP() represents the transposed convolution, which is used for image upsampling operation, and PRSSB is the progressive state space module.
[0033] A further improvement of the present invention is that the progressive state space module PRSSB includes a variational state space model VSSM, a channel attention layer, an enhanced spatial attention layer, and a convolution layer with a convolution kernel size of 3. The specific method of the progressive state space module PRSSB is as follows: for a given output feature F in :
[0034] F VSSM1 =S(L(F in )
[0035] F VSSM2 =LN(2D(S(C3(L(F in )))))
[0036] F VSSM =L(F VSSM1 ⊙F VSSM2 )
[0037] Among them, F VSSM1 is the output result of the upper branch of the variational state space model VSSM, F VSSM2 is the output result of the branch under the variational state space model VSSM, S() represents the SiLU activation function, L() represents the fully connected layer, LN() represents the normalization operation, 2D() is the 2D scanning module, C3 represents the convolution layer with a convolution kernel size of 3, and F VSSM It is the output feature of the VSSM part on the left side of the progressive state space module PRSSB, and then obtains information through channel attention and enhanced spatial attention:
[0038] F CA =CA(C3(F VSSM ))
[0039] F ESA =ESA(C3(F VSSM ))
[0040]
[0041] Among them, CA() represents channel attention, C3 represents the convolution layer with a convolution kernel size of 3, ESA represents enhanced spatial attention, and F CA and F ESA are the output results of the feature after channel attention and output attention, Concat() is the splicing operation according to the channel dimension, F out is the output feature of the progressive state space module PRSSB.
[0042] A further improvement of the present invention is that the reconstruction module converts the output features F of the four branches in the deep feature extraction module into D1 、F D2 、F D3 、F D4 Through the channel shuffling strategy, the final reconstruction is achieved using the sub-pixel upsampling module to obtain the final high-resolution reconstructed image I SR , the process is described as:
[0043] I SR =H PixeShuffle (C3(F final ))
[0044] F final =C1(CShuffle(Concat(Concat(F D1 ,F D2 ),Concat(F D3 ,F D4 ))))
[0045] Among them, C1 represents the convolution layer with a convolution kernel size of 1 for dimensionality reduction operation, CShuffle() represents the channel shuffling strategy, Concat() is the splicing operation according to the channel dimension, and F final is the output feature after multi-branch fusion, H PixeShuffle (·) is sub-pixel convolution.
[0046] A further improvement of the present invention is that the super-resolution reconstruction system includes a camera, a Linux platform, a Windows platform, a display and a buzzer. The camera obtains on-site images. The Linux platform and the Windows platform are used to deploy a super-resolution reconstruction method for blurred images in a deep-earth disaster environment. The display displays the violation picture. The buzzer is used for voice reminders. The super-resolution reconstruction method for blurred images in a deep-earth disaster environment obtains the original image of the mine from the camera through the real-time streaming protocol RTSP, and sends the generated high-resolution mine image to the violation detection model. The violation detection model is a violation detection model based on yolov5 in the prior art. The detection accuracy is improved by enhancing the resolution of the input image. In the event of a violation, the super-resolution reconstruction method for blurred images in a deep-earth disaster environment realizes a voice alarm by controlling the output GPIO port of the buzzer, and realizes an image alarm by transmitting it to the display through websocket.
[0047] The beneficial effects of the present invention are:
[0048] This paper, based on a pyramid structure combined with the Mamba model, fully considers the importance of a large receptive field in image reconstruction tasks, effectively capturing global information within an image. This structure successfully achieves a good balance between the number of model parameters, computational complexity, and reconstruction performance, enabling the network to be more efficient in practical applications while maintaining high performance.
[0049] This paper combines the pyramid structure of convolutional neural networks with state-space models to improve image processing and analysis performance. Specifically, the pyramid structure effectively captures information at different levels within an image through multi-scale feature extraction, thereby improving the model's ability to understand complex visual tasks. The state-space model, on the other hand, provides a dynamic modeling framework, enabling the model to better capture state changes and feature evolution in continuous data.
[0050] Taking into account that shallow features have a good guiding role in image reconstruction, the present invention designs a shallow feature extraction method of a progressive two-way convolution chain, which obtains the overall structural information and representative image features of the image by gradually adjusting the convolution kernel size. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is a flow chart of the reconstruction method of the present invention.
[0052] Figure 2 It is a structural diagram of the reconstruction method of the present invention.
[0053] Figure 3 It is a structural diagram of the shallow feature extraction module of the present invention.
[0054] Figure 4It is a structural diagram of the progressive state space module PRSSB of the present invention.
[0055] Figure 5 Schematic diagram of the reconstruction system of the present invention.
[0056] Figure 6 Schematic diagram of the reconstruction process of the image reconstruction module of the present invention. DETAILED DESCRIPTION
[0057] The following drawings illustrate embodiments of the present invention. For clarity, many practical details are included in the following description. However, it should be understood that these practical details are not intended to limit the present invention. In other words, in some embodiments of the present invention, these practical details are not essential. Furthermore, to simplify the drawings, some commonly used structures and components are depicted in simplified schematic form.
[0058] like Figure 5 As shown, this paper proposes a super-resolution reconstruction system for blurred images in deep-earth disaster environments. This system aims to restore low-quality blurred images in disaster environments to high-definition images through efficient information processing techniques. This improves the accuracy of subsequent emergency rescue, violation detection, and identification tasks, provides more precise visual support for post-disaster recovery and rescue decision-making, and promotes the development and application of intelligent rescue systems in deep-earth disaster environments. The system can be deployed on multiple platforms, including Linux and Windows, and uses the RTSP protocol to obtain low-resolution original images. Specifically, the super-resolution reconstruction system includes a camera, a Linux platform, a Windows platform, a display and a buzzer. The camera obtains on-site images. The Linux platform and the Windows platform are used to deploy a super-resolution reconstruction method for blurred images in a deep-ground disaster environment. The display displays the violation picture. The buzzer is used for voice reminders. The super-resolution reconstruction method for blurred images in a deep-ground disaster environment obtains the original image of the mine from the camera through the real-time streaming protocol RTSP, and sends the generated high-resolution mine image to the violation detection model. The violation detection model is a violation detection model based on yolov5 in the existing technology. The detection accuracy is improved by enhancing the resolution of the input image. In the event of a violation, the super-resolution reconstruction method for blurred images in a deep-ground disaster environment realizes a voice alarm by controlling the output GPIO port of the buzzer, and realizes an image alarm by transmitting it to the display through websocket.
[0059] The present invention provides a method for super-resolution reconstruction of images in a blurred field of view in a deep earth disaster environment. The method is implemented by a progressive Mamba network. The progressive Mamba network includes a shallow feature extraction module composed of two progressive convolution chains with opposite convolution orders, a deep feature extraction module, and a reconstruction module composed of a channel shuffling strategy and a sub-pixel upsampling module.
[0060] like Figure 3 As shown, the shallow feature extraction module includes two convolution layers with a convolution kernel size of 5, two convolution layers with a convolution kernel size of 3, and three convolution layers with a convolution kernel size of 1. The two progressive convolution chains with opposite convolution orders in the shallow feature extraction module are the upper branch convolution chain and the lower branch convolution chain. The upper branch convolution chain and the lower branch convolution chain respectively include a convolution layer with a convolution kernel size of 5, a convolution layer with a convolution kernel size of 3, and a convolution layer with a convolution kernel size of 1. In the upper branch convolution chain, the low-resolution image I LR First, it passes through a convolution layer with a convolution kernel size of 5 to enlarge the receptive field and extract more representative features. Then, it passes through a convolution layer with a convolution kernel size of 3 and a convolution layer with a convolution kernel size of 1 to refine the low-resolution image I. LR Features, while retaining global information, pay more attention to the extraction of local features, enhancing the richness and diversity of features. The lower branch convolution chain is opposite, the low resolution image I LR First, it passes through a convolution layer with a convolution kernel size of 1, and then passes through a convolution layer with a convolution kernel size of 3 and a convolution layer with a convolution kernel size of 5 in sequence, gradually increasing the receptive field and effectively integrating more contextual information, thereby improving the understanding of the overall structure of the image. At the same time, the output features of the upper branch convolution chain and the lower branch convolution chain are spliced together according to the channel dimension, and after channel shuffling, a convolution layer with a convolution kernel size of 1 is used for dimensionality reduction to obtain the output feature F containing local detail features and global overall structure information. S .
[0061] like Figure 2As shown, the first branch in the deep feature extraction module of the present invention includes a convolution layer with a convolution kernel size of 5, a convolution layer with a convolution kernel size of 3, and a convolution layer with a convolution kernel size of 1. In the first branch, a convolution chain that gradually adjusts the convolution kernel size is adopted to achieve accurate extraction of image features. At the same time, through multiple upsampling operations, the first branch can effectively increase the image size, thereby retaining more detail information and contextual features. The second branch in the deep feature extraction module includes three progressive state space modules PRSSB, the third branch in the deep feature extraction module includes a convolution layer with a convolution kernel size of 5, an upsampling operation, and a convolution layer with a convolution kernel size of 3. The fourth branch in the deep feature extraction module is a convolution layer with a convolution kernel size of 5. Among them, as Figure 4 As shown in the figure, the progressive state space module PRSSB includes a variational state space model VSSM, a channel attention layer, an enhanced spatial attention layer, and a convolutional layer with a convolution kernel size of 3.
[0062] like Figure 1 As shown, the method for super-resolution reconstruction of images in a blurred field of view in a deep earth disaster environment of the present invention specifically includes the following steps:
[0063] Step 1: Take the high-resolution image I HR Perform bicubic downsampling to obtain a low-resolution image I LR , the low-resolution image I is obtained by the shallow feature extraction module LR Perform shallow feature extraction. The extraction process is described as follows:
[0064] F S =H Shallow (I LR )
[0065] Among them, H Shallow (·) is the shallow feature extraction module, F S It is the output feature of the shallow feature extraction module.
[0066] Among them, the upper branch convolution chain and the lower branch convolution chain are used to process the low-resolution image I LR The processing process is:
[0067] F S1 =C1(C3(C5(I LR )));
[0068] F S2 =C5(C3(C1(I LR )));
[0069] F S =C1(CShuffle(Concat(F S1 ,FS2 )))
[0070] Among them, C5 represents the convolution layer with a convolution kernel size of 5, C3 represents the convolution layer with a convolution kernel size of 3, C1 represents the convolution layer with a convolution kernel size of 1, and F S1 is the output result of the upper branch convolution chain, F S2 is the output result of the convolution chain of the lower branch, Concat() represents splicing by channel dimension, CShuffle() represents the channel shuffling strategy, F S It is the output feature of the shallow feature extraction module.
[0071] Step 2: Input the output features of the shallow feature extraction module into the deep feature extraction module to extract deep features, wherein the deep feature extraction module includes four branches of three progressive state space modules PRSSB, each branch processes the output features of the input shallow feature extraction module respectively, and outputs an output feature respectively.
[0072] In this step, the output features of the input shallow feature extraction module are processed by the deep feature extraction module. The specific method of extracting deep features is:
[0073] Step 2.1: Extract the output feature F of the shallow feature extraction module from the input through the first branch S Processing is performed to obtain the output feature F of the first branch D1 :
[0074] F D1 =C1(UP(C3(UP(C5(F S )))))
[0075] Step 2.2: The output feature F of the shallow feature extraction module of the input is extracted through the second branch S After upsampling and processing by the progressive state space module PRSSB, the output feature F of the second branch is obtained D2 :
[0076] F D2 =PRSSB(UP(PRSSB(UP(PRSSB(F S )))))
[0077] Step 2.3: Upsample the output of the first progressive state space module PRSSB in the second branch through the third molecule to obtain the output feature F of the third branch. D3 :
[0078] F D3 =C3(UP(C5(UP(PRSSB(F S )))))
[0079] Step 2.4: After upsampling the output result of the second progressive state space module PRSSB in the second branch through the fourth branch, the output feature F of the fourth branch is obtained. D4 :
[0080] F D4 =C5(UP(PRSSB(UP(PRSSB(F S )))))
[0081] Among them, F S is the output feature of the shallow feature extraction module, C5 represents the convolution layer with a convolution kernel size of 5, C3 represents the convolution layer with a convolution kernel size of 3, C1 represents the convolution layer with a convolution kernel size of 1, UP() represents the transposed convolution, which is used for image upsampling operation, and PRSSB is the progressive state space module.
[0082] like Figure 4 As shown, for the convenience of description, the present invention gives the output feature F in , for the first progressive state space module PRSSB, the output feature F in The output feature F of the extracted shallow feature extraction module S ,For the second PRSSB, output feature F in is the output feature of the first PRSSB, and for the third PRSSB, the output feature F in is the output feature of the second PRSSB. Therefore, the specific method for the progressive state space module PRSSB to process the feature is:
[0083] F VSSM1 =S(L(F in )
[0084] F VSSM2 =LN(2D(S(C3(L(F in )))))
[0085] F VSSM =L(F VSSM1 ⊙F VSSM2 )
[0086] Among them, F VSSM1 is the output result of the upper branch of the variational state space model VSSM, F VSSM2 is the output result of the branch under the variational state space model VSSM, S() represents the SiLU activation function, L() represents the fully connected layer, LN() represents the normalization operation, 2D() is the 2D scanning module, C3 represents the convolution layer with a convolution kernel size of 3, and F VSSMIt is the output feature of the VSSM part on the left side of the progressive state space module PRSSB, and then obtains information through channel attention and enhanced spatial attention:
[0087] F CA =CA(C3(F VSSM ))
[0088] F ESA =ESA(C3(F VSSM ))
[0089]
[0090] Among them, CA() represents channel attention, C3 represents the convolution layer with a convolution kernel size of 3, ESA represents enhanced spatial attention, and F CA and F ESA are the output results of the feature after channel attention and output attention, Concat() is the splicing operation according to the channel dimension, F out is the output feature of the progressive state space module PRSSB.
[0091] Step 3: Fuse the output features of the four branches in step 2 together, and use the reconstruction module to reconstruct the fused features to obtain the final high-resolution reconstructed image I SR .
[0092] The reconstruction module converts the output features F of the four branches in the deep feature extraction module into D1 、F D2 、F D3 、F D4 Through the channel shuffling strategy, the final reconstruction is achieved using the sub-pixel upsampling module to obtain the final high-resolution reconstructed image I SR , the process is described as:
[0093] I SR =H PixeShuffle (C3(F final ))
[0094] F final =C1(CShuffle(Concat(Concat(F D1 ,F D2 ),Concat(F D3 ,F D4 ))))
[0095] Among them, C1 represents the convolution layer with a convolution kernel size of 1 for dimensionality reduction operation, CShuffle() represents the channel shuffling strategy, Concat() is the splicing operation according to the channel dimension, and F finalis the output feature after multi-branch fusion, H PixeShuffle (·) is sub-pixel convolution.
[0096] The basic principle of sub-pixel convolution is to divide the input low-resolution feature image into several non-overlapping pixel blocks, and then expand these pixel blocks to the high-resolution target image size through convolution. During the convolution operation, the dimensions of each pixel block are expanded and locally connected with neighboring pixel blocks. In other words, the low-frequency information originally in the feature map is spatially distributed through convolution to generate a high-resolution output.
[0097] For the feature F fed into the sub-pixel convolution final , whose tensor dimensions are H×W×C. Sub-pixel convolution first uses a kernel×kernal×C×r 2 C standard convolution layer to F final After preliminary processing, the dimension of the output feature map is H×W×r 2 C. Then, sub-pixel convolution divides the feature map into channels and interleaves the slices along the channel dimension to achieve the rearrangement of each feature point. The size of the rearranged feature map is rH×rW×C, where each feature point contains information from r×r pixel blocks. Finally, a standard convolution operation is performed on the rearranged feature map to reduce the number of channels to 3 (color image), which gives the final high-resolution image, as shown in Figure 6 shown.
[0098] This paper fully considers the importance of a large receptive field in image reconstruction tasks and can effectively capture global information in the image. By combining the pyramid structure with the Mamba model, it achieves a good balance between the number of model parameters, computational complexity, and reconstruction performance.
[0099] The foregoing is merely an embodiment of the present invention and is not intended to limit the present invention. It will be apparent to those skilled in the art that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.
Claims
1. A method for super-resolution reconstruction of images in a deep-earth disaster environment with a blurred field of view, characterized by: The method for super-resolution reconstruction of images in a blurred field of view in a deep earth disaster environment is implemented by a progressive Mamba network. The progressive Mamba network includes a shallow feature extraction module composed of two progressive convolution chains with opposite convolution orders, a deep feature extraction module, and a reconstruction module composed of a channel shuffling strategy and a sub-pixel upsampling module. The specific reconstruction method includes the following steps: Step 1: Take the high-resolution image captured by the camera Perform bicubic downsampling to obtain a low-resolution image , the low-resolution image obtained by the shallow feature extraction module Perform shallow feature extraction. The extraction process is described as follows: , in, is a shallow feature extraction module, Output features of the shallow feature extraction module; Step 2: Input the output features of the shallow feature extraction module into the deep feature extraction module to extract deep features, wherein the deep feature extraction module includes three progressive state space modules There are four branches, each of which processes the output features of the input shallow feature extraction module and outputs an output feature respectively; Step 3: Fuse the output features of the four branches in step 2 together, and use the reconstruction module to reconstruct the fused features to obtain the final high-resolution reconstructed image. ,in: The output features of the input shallow feature extraction module are processed by the deep feature extraction module. The specific method of extracting deep features is as follows: Step 2.1: Extract the output features of the shallow feature extraction module from the input through the first branch Processing is performed to obtain the output features of the first branch : , Step 2.2: Extract the output features of the shallow feature extraction module from the input through the second branch After upsampling and progressive state-space modules Processing is performed to obtain the output features of the second branch : , Step 2.3: The first progressive state space module in the second branch is applied through the third molecule The output result is upsampled and processed to obtain the output features of the third branch : , Step 2.4: Use the fourth branch to train the second progressive state space module in the second branch. After upsampling the output result, the output feature of the fourth branch is obtained. : , in, is the output feature of the shallow feature extraction module, represents a convolution layer with a convolution kernel size of 5, represents a convolution layer with a convolution kernel size of 3, represents a convolution layer with a convolution kernel size of 1, Represents transposed convolution, used for image upsampling operation, It is a progressive state-space module.
2. The method for super-resolution reconstruction of images in a deep earth disaster environment with a blurred field of view according to claim 1, characterized in that: The two progressive convolution chains with opposite convolution orders in the shallow feature extraction module are an upper branch convolution chain and a lower branch convolution chain, and the upper branch convolution chain and the lower branch convolution chain respectively include a convolution layer with a convolution kernel size of 5, a convolution layer with a convolution kernel size of 3, and a convolution layer with a convolution kernel size of 1.
3. The method for super-resolution reconstruction of images in a blurred field of view in a deep earth disaster environment according to claim 2, characterized in that: The upper branch convolution chain and the lower branch convolution chain are used to perform low-resolution image The processing process is: ; ; , in, represents a convolution layer with a convolution kernel size of 5, represents a convolution layer with a convolution kernel size of 3, represents a convolution layer with a convolution kernel size of 1, is the output result of the upper branch convolution chain, is the output result of the next branch convolution chain, Indicates splicing by channel dimension, represents the channel shuffling strategy, It is the output feature of the shallow feature extraction module.
4. The method for super-resolution reconstruction of images in a blurred field of view in a deep earth disaster environment according to claim 1, characterized in that: The progressive state-space module Including variational state space models , channel attention layer, enhanced spatial attention layer, convolution layer with kernel size 3, progressive state space module The specific method of processing features is: for a given output feature : , in, is a variational state space model The output of the upper branch is: is a variational state space model The next branch outputs the result. represents the SiLU activation function, represents the fully connected layer, Indicates the normalization operation. For the 2D scanning module, represents a convolution layer with a convolution kernel size of 3, For the progressive state space module Left side Part of the output features, followed by channel attention and enhanced spatial attention to obtain information: , in, represents channel attention, represents a convolution layer with a convolution kernel size of 3, Indicates enhanced spatial attention, and are the output results of feature channel attention and output attention respectively, is the splicing operation according to the channel dimension, For the progressive state space module The output features of .
5. The method for super-resolution reconstruction of images in a blurred field of view in a deep earth disaster environment according to claim 4, characterized in that: The reconstruction module extracts the output features of the four branches in the deep feature extraction module. 、 、 、 Through the channel shuffling strategy, the final reconstruction is achieved using the sub-pixel upsampling module to obtain the final high-resolution reconstructed image. , the process is described as: , , in, represents a convolutional layer with a convolution kernel size of 1 for dimensionality reduction operations, represents the channel shuffling strategy, is the splicing operation according to the channel dimension, is the output feature after multi-branch fusion, It is sub-pixel convolution.
6. A super-resolution reconstruction system for blurred images in a deep-earth disaster environment implementing the reconstruction method according to any one of claims 1 to 5, characterized in that: The super-resolution reconstruction system includes a camera, a Linux platform, a Windows platform, a display and a buzzer. The camera obtains on-site images. The Linux platform and the Windows platform are used to deploy a super-resolution reconstruction method for blurred images in a deep-ground disaster environment. The display displays the violation picture. The buzzer is used for voice reminders. The super-resolution reconstruction method for blurred images in a deep-ground disaster environment obtains the original image of the mine from the camera through the real-time streaming protocol RTSP, and sends the generated high-resolution mine image to the violation behavior detection model. The detection accuracy is improved by enhancing the input image resolution. In the event of a violation, the super-resolution reconstruction method for blurred images in a deep-ground disaster environment realizes a voice alarm by controlling the output GPIO port of the buzzer, and realizes an image alarm by transmitting it to the display through websocket.
Citation Information
Patent Citations
Image restoration method based on state space model
CN118195905A
Lightweight mine image super-resolution reconstruction system and method based on progressive receptive field
CN118918005A