Image rain removal method based on multi-scale state space model
Through the image rain removal method based on the multi-scale state space model, the multi-scale Mamba block and frequency feature enhancement module of the multi-scale pyramid image and encoder-decoder network are used to solve the problems of high computing costs and difficulty in capturing global features in the prior art, and the efficient and low-complexity image rain removal effect is achieved.
Patent Information
- Application Number
- CN202510510624.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-23
AI Technical Summary
Existing image rain removal methods are cost-effective when processing high-resolution images and are difficult to effectively capture global features and multi-scale information.
The image rain removal method based on the multi-scale state space model is adopted, and the image is decomposed into a multi-scale pyramid image through downsampling, and the multi-scale Mamba block and frequency feature enhancement module in the encoder-decoder network are used for feature extraction and image reconstruction, and finally the rain removal image is obtained by addition.
It significantly reduces the computational complexity and can effectively capture global features and multi-scale information, thereby achieving high-quality image de-rain, suitable for deployment on resource-constrained mobile or embedded devices.
Smart Images

Figure CN120031750A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and image processing, and in particular to an image deraining method based on a multi-scale state space model. Background Art
[0002] As a common weather phenomenon, rain not only significantly reduces the visibility of images, but also may obscure key visual details, posing a major challenge to downstream visual tasks in application areas such as autonomous driving and video surveillance. The goal of image deraining is to remove the undesirable degradation caused by rain from the input image, thereby improving its visual quality and improving the accuracy of the perception system. Therefore, conducting effective image deraining research can not only improve the visual quality of images, but also has important research significance.
[0003] Many methods have been proposed for image deraining. Early methods based on prior knowledge usually rely on model-driven strategies, using prior information of clean images or specific features of rain streaks to guide the restoration process. Due to the ill-posed nature of image deraining, traditional methods usually rely on image decomposition, low-rank representation, discriminative sparse coding, and Gaussian mixture models. However, these methods often rely on empirical observations and lead to complex optimization problems.
[0004] Subsequently, the rise of deep learning has greatly promoted the development of image deraining. Convolutional neural network-based methods perform well in capturing the complex mapping relationship between rainy images and clear images, enabling them to effectively process rain streaks of different shapes, sizes, and densities. However, since convolutional neural networks have a fixed receptive field, this limits their ability to capture global features and multi-scale information.
[0005] Zamir et al. designed an efficient Transformer model by estimating self-attention along the channel dimension and achieved remarkable performance. Xiao et al. first introduced the Transformer for image deraining based on spatial and window self-attention modules. Chen et al. developed a sparse Transformer to make full use of the most useful features for better image restoration. However, the self-attention mechanism introduces quadratic spatial and temporal complexity, which significantly increases the computational cost when processing high-resolution images, which is unacceptable. Therefore, it is necessary to explore a method that does not significantly increase the computational cost and can effectively capture non-local information to achieve high-quality image deraining. Summary of the invention
[0006] The purpose of the embodiment of the present invention is to provide an image deraining method based on a multi-scale state space model. To achieve the above object, an embodiment of the present invention provides an image de-raining method based on a multi-scale state space model, including: Obtain an image with rain streaks, and decompose the image with rain streaks into a multi-scale pyramid image by downsampling; Input the image with rain streaks and the multi-scale pyramid image into a pre-constructed multi-scale state space model respectively. After feature extraction and image reconstruction through a convolutional layer and an encoder-decoder network, a reconstructed residual image is obtained, where the encoder-decoder network includes a multi-scale Mamba block and a frequency feature enhancement module; Add the reconstructed residual image to the image with rain streaks to obtain a target de-rained image.
[0007] Optionally, the multi-scale Mamba block is used to perform a geometric transformation operation on the multi-scale pyramid image by combining a multi-scale 2D selective scanning mechanism, and generate scanning sequences in different directions to represent global feature information; The frequency feature enhancement module is used to map the input features to the frequency domain through Fourier transform, separate the real part and the imaginary part and splice them into the channel dimension, extract frequency domain features through convolution and non-linear activation, and restore them to the spatial domain through inverse Fourier transform to achieve the extraction of deep features.
[0008] Optionally, the multi-scale 2D selective scanning mechanism is represented by the following formula:
[0009] In the formula, k represents the scale size, k = 1 represents the small scale, k = 2 represents the medium scale, k = 4 represents the large scale, X represents the input image, Stack() represents the stacking operation of the image, T() represents the transpose operation of the image, F() represents the flipping operation of the pixel, and Cat() represents the splicing of the images.
[0010] Optionally, the encoder-decoder network further includes a gated feature fusion module for adaptively aggregating complementary features of multi-scale branches.
[0011] Optionally, inputting the image with rain streaks and the multi-scale pyramid image into a pre-constructed multi-scale state space model respectively, and obtaining a reconstructed residual image after feature extraction and image reconstruction through a convolutional layer and an encoder-decoder network includes: Input the image with rain streaks and the multi-scale pyramid image into the convolutional layer of the pre-constructed multi-scale state space model respectively to extract shallow features; The shallow features are input into the multi-scale Mamba block in the encoder-decoder network of each scale branch, and two parallel branches are processed, and the features of the two branches are aggregated and output through the Hadamard product; wherein, in the first branch, the feature channel is processed by deep convolution and SiLU activation function, and combined with the multi-scale 2D selective scanning strategy and layer normalization operation, and in the second branch, layer normalization and SiLU activation operations are performed; The features of different scales are aggregated through the gated feature fusion module to obtain the aggregated features; The aggregated features are input into the decoder network for image reconstruction to obtain a reconstructed residual image; wherein the multi-scale state space model network architecture includes three scale branches, each scale branch includes an encoder and a decoder.
[0012] Optionally, features of different scales are aggregated through a gated feature fusion module, and the aggregated features are obtained according to the following formula: ; ; In the formula, Input features for different scales, express Activation function, represents 3×3 convolution, represents 1×1 convolution, ⊙ represents element-wise multiplication, represents the input features, represents the features after aggregation, Represents a gating unit.
[0013] Optionally, the loss function used by the multi-scale state-space model is as follows:
[0014] In the formula, represents the Charbonnier loss, represents the edge loss, represents the frequency loss, , and Represent weights respectively.
[0015] In a second aspect, the present invention provides an image deraining system based on a multi-scale state space model, comprising: An acquisition unit, used for acquiring an image with rain streaks, and decomposing the image with rain streaks into a multi-scale pyramid image by downsampling; A reconstruction unit, used for inputting the image with rain streaks and the multi-scale pyramid image into a pre-constructed multi-scale state space model, respectively, and obtaining a reconstructed residual image after performing feature extraction and image reconstruction through a convolutional layer and an encoder-decoder network, wherein the encoder-decoder network includes a multi-scale Mamba block and a frequency feature enhancement module; The obtaining unit is used to add the reconstructed residual image and the image with rain streaks to obtain a target rain-free image.
[0016] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned image deraining method based on the multi-scale state space model when executing the program.
[0017] In a fourth aspect, the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-mentioned image deraining method based on a multi-scale state space model.
[0018] Through the above technical solution, the state space model is used to model the global information. At the same time, the multi-scale complementary information is effectively utilized in combination with the multi-scale framework, and the cross-scale complementary features are explicitly mined to more accurately remove rain streaks. High-quality clear images are gradually restored from rainy images, ensuring the clarity and visual quality of the derained images.
[0019] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following specific implementations, they are used to explain the embodiments of the present invention, but do not constitute a limitation on the embodiments of the present invention. In the accompanying drawings: Figure 1 is an implementation flow chart of an image deraining method based on a multi-scale state space model provided by an embodiment of the present invention; Figure 2 is a schematic diagram of a multi-scale 2D selective scanning strategy provided by an embodiment of the present invention; Figure 3 It is a schematic diagram of a frequency feature enhancement module and a gated feature fusion module provided by an embodiment of the present invention; Figure 4 Schematic diagram of the overall network structure of an image deraining method based on a multi-scale state space model provided by an embodiment of the present invention; Figure 5It is a structural schematic diagram of an image rain removal system based on a multi-scale state space model provided by an embodiment of the present invention; Figure 6 The present invention is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0021] The specific implementation of the embodiment of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the embodiment of the present invention, and is not used to limit the embodiment of the present invention.
[0022] See also Figure 1 As shown, it is a flowchart of an implementation of an image deraining method based on a multi-scale state space model provided by an embodiment of the present invention, including the following execution steps: Step 100: Acquire an image with rain streaks, and decompose the image with rain streaks into multi-scale pyramid images by downsampling.
[0023] Specifically, the input image is down-sampled to generate three scale branches of the pyramid scale, namely, the large scale, i.e., the original image, the middle scale, i.e., the 1 / 2 sampling branch, and the small scale, i.e., the 1 / 4 sampling branch.
[0024] Step 101: respectively input the image with rain streaks and the multi-scale pyramid image into a pre-constructed multi-scale state space model, perform feature extraction and image reconstruction through a convolutional layer and an encoder-decoder network, and obtain a reconstructed residual image.
[0025] The encoder-decoder network includes a multi-scale Mamba block and a frequency feature enhancement module.
[0026] Specifically, the multi-scale Mamba block is used to combine the multi-scale 2D selective scanning mechanism to apply geometric transformation operations to the multi-scale pyramid image, and generate scanning sequences in different directions to represent global feature information.
[0027] Furthermore, the multi-scale Mamba block models global features through a state space model, as follows: The state space model is a mathematical framework commonly used in time series analysis and control systems. The state equation describes the evolution of the underlying system over time and characterizes the relationship between the hidden state of the system and its time dynamics. The input signal x(t) ∈ R is represented by the implicit latent state Mapped to the output response y(t) ∈ R. It can be expressed as a set of first-order linear ordinary differential equations as follows: h′(t) = Ah(t) + Bx(t) y(t) = Ch(t) + Dx(t) Where N represents the state size, and A, B, C, and D are learnable weight matrices.
[0028] Subsequently, the above equations are usually integrated into the actual deep learning algorithm by using a discretization process. Specifically, let Δ be the time scale parameter used to convert the continuous parameters A and B into discrete parameters , The commonly used discretization method is the zero-order hold rule, which is expressed as follows:
[0029]
[0030]
[0031] However, the above formula is mainly for linear time-invariant systems whose parameters remain unchanged under different inputs. To overcome this limitation, Mamba improves the state-space module and proposes a selective scanning mechanism (S6) to improve the original state-space module by introducing specific restoration prior knowledge to simultaneously achieve input-dependent weights and linear computational complexity.
[0032] Specifically, local patch repetitiveness and inter-channel interactions are taken into account to assist long-range spatial modeling in Mamba. Given the input deep features , where H and W represent the height and width of the image, C represents the channel, first layer normalization (LN) is used, and then the visual state space block (VSSB) is used to capture spatial long-range dependencies. In order to improve network performance, a learnable scaling factor is used in the residual connection , the specific expression is:
[0033] Then, another normalization layer (LN) is used to Normalization is performed and a convolutional layer is used to model the spatial local similarity prior. Then a forward propagation layer (FFL) is used to obtain the final output of VSSB. , which can be expressed as:
[0034] For example, in a multi-scale network architecture, different scale branches contain different rain degradation features. Compared with small-scale branches, large-scale branches contain richer feature information. If the same scanning strategy is used for branches of different scales, it may lead to information redundancy and waste of computing resources. Figure 2 As shown in FIG, the multi-scale Mamba block adopts a multi-scale 2D selective scanning strategy to better extract the potential explicit information in different scales. Figure 2 As shown in Figure 1, the module assigns different numbers of scanning directions to branches of different scales through geometric transformation, thereby achieving efficient sequential scanning. Specifically, larger scales use more scanning directions (four scanning directions), medium scales are assigned one scanning direction, and smaller scales use fewer scanning directions. The multi-scale 2D selective scanning mechanism is expressed according to the following formula:
[0035] Where k represents the scale, k = 1 represents a small scale, k = 2 represents a medium scale, k = 4 represents a large scale, X represents the input image, Stack() represents the stacking operation of the image, T() represents the transposition operation of the image, F() represents the flipping operation of the pixel, and Cat() represents the splicing of the image.
[0036] Specifically, the frequency feature enhancement module is used to map the input features to the frequency domain through Fourier transform, separate the real part and the imaginary part and splice them into channel dimensions, extract the frequency domain features through convolution and nonlinear activation, and restore them to the spatial domain through inverse Fourier transform to achieve deep feature extraction.
[0037] Furthermore, the frequency characteristic enhancement module is as follows: Figure 3 As shown, the frequency feature enhancement module is specifically: input It is divided into two branches. One branch extracts local features in the spatial domain through time domain convolution operation, while the other branch uses two-dimensional real fast Fourier transform to map the input features to the frequency domain to obtain F(X)∈R H ×W / 2×C Subsequently, the real and imaginary features of F(X) are concatenated in the channel dimension to obtain Y∈RH× W / 2 ×2C, and then the processed features are restored to the time domain through inverse Fourier transform to obtain the output Finally, through residual connection, the spatial domain features, frequency domain features and original input are added to obtain the final output.
[0038] Preferably, the encoder-decoder network further comprises a gated feature fusion module for adaptively aggregating complementary features of multi-scale branches.
[0039] Specifically, when executing step 101, the following steps may be specifically performed: S1010: respectively inputting the image with rain streaks and the multi-scale pyramid image into the convolution layer of a pre-constructed multi-scale state space model to extract shallow features.
[0040] S1011: Input the shallow features into the multi-scale Mamba block in the encoder-decoder network of each scale branch, perform two parallel branch processing, and aggregate and output the features of the two branches through the Hadamard product.
[0041] In the first branch, the feature channel is processed by deep convolution and SiLU activation function, and combined with multi-scale 2D selective scanning strategy and layer normalization operation. In the second branch, layer normalization and SiLU activation operations are performed.
[0042] In this embodiment, the multi-scale Mamba block is specifically implemented as follows: input feature The processing is done in two parallel branches. In the first branch, the feature channel is processed by deep convolution and SiLU activation function, combined with multi-scale 2D selective scanning strategy and layer normalization operation. In the second branch, layer normalization and SiLU activation operation are performed. After that, the features of the two branches are aggregated by Hadamard product. The final output is , the specific process can be expressed as:
[0043] Where DWConv represents depthwise separable convolution and ⊙ represents the Hadamard product.
[0044] S1012: Aggregate features of different scales through a gated feature fusion module to obtain aggregated features.
[0045] For details, see the gated feature fusion module. Figure 3 As shown in FIG. 1 , the gated feature fusion module is designed to better promote the fusion of features at different scales. For input features from three scales, a dual-input gated unit is used, which can dynamically adapt to inputs at different scales. The formula of the gated unit can be expressed as: ; in, represents 3×3 convolution, represents 1×1 convolution, Tanh() represents tanh activation function, and ⊙ represents element-by-element multiplication. By combining these two gating units, different features can be effectively fused, and the process can be expressed as: ; In the formula, represents the input features, represents the features after aggregation, Represents a gating unit.
[0046] Furthermore, the aggregated features of small and medium scales are spliced to the previous scale through the decoder network composed of multi-scale Mamba blocks and frequency feature enhancement modules, and finally fused into the large-scale network for decoding. Each scale is then reconstructed through 3×3 convolution to obtain the residual image. .
[0047] S1013: Input the aggregated features into the decoder network for image reconstruction to obtain a reconstructed residual image.
[0048] The multi-scale state space model network architecture includes three scale branches, each of which includes an encoder and a decoder. Each branch uses a multi-scale Mamba module combined with a multi-scale 2D selective scanning strategy to capture global feature information. Then, the frequency feature enhancement module extracts local feature information in the frequency domain, which helps to better promote image restoration. In addition, the complementary features between scales are adaptively aggregated through the gated feature fusion module between scales, further improving the high-quality restoration effect of the image and achieving efficient deraining performance.
[0049] More specifically, during the training process of the model, in order to supervise the training process of the network and enhance its ability to capture details, a weighted sum of three losses is selected as the basis for guiding network training. The loss function used by the multi-scale state space model is as follows:
[0050] In the formula, represents the Charbonnier loss, represents the edge loss, represents the frequency loss, , and Represent weights respectively.
[0051] For example, the scalar weight , and They are empirically set to 1, 0.05, and 0.01, respectively.
[0052] Step 102: Add the reconstructed residual image to the image with rain streaks to obtain a target rain-free image.
[0053] Furthermore, the reconstructed residual image is added to the original rainy image to obtain the final rain-free image: .
[0054] In this embodiment, the overall network structure diagram of the image deraining method based on the multi-scale state space model is shown in Figure 4As shown in Figure 2, the multi-scale state space model network architecture contains three scale branches, each of which contains an encoder and a decoder. For a given rainy day input image ,First, the coarse-to-fine technique is used to divide it into three scales, that is, the input image is interpolated by the Downsample to 1 / 2 and 1 / 4 scale to generate pyramid-shaped multi-scale rainy day images. These images are used as inputs of each scale and shallow features are extracted through 3×3 convolutional layers. , where H×W represents the spatial dimension and C is the number of feature channels. Next, the shallow features are input into the encoder-decoder network of each scale branch. The encoder / decoder of each branch consists of a multi-scale Mamba block and a frequency feature enhancement module. The multi-scale Mamba block combines a multi-scale 2D selective scanning mechanism to capture global feature information with linear computational complexity, while the frequency feature enhancement module focuses on local feature information through the frequency domain to achieve deep feature extraction. In addition, in this network, a gated feature fusion module is introduced to aggregate feature selection information between different scales and representations of a specific scale, and input it into the decoder network to reconstruct a high-quality output image to obtain a reconstructed residual image. Finally, the reconstructed residual image is added to the original rainy image to obtain the final derained image. .
[0055] Furthermore, in order to verify the effectiveness of this application, the deraining effect of the model is evaluated on common image deraining public benchmark datasets such as Rain200H. On the above datasets, the peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) of the Y channel in the YCbCr color space are calculated as benchmark evaluation indicators for quantitative comparison.
[0056] In this embodiment, the peak signal-to-noise ratio (PSNR) is used to measure the difference between two images, such as the difference between a compressed image and an original image, to evaluate the quality of the compressed image; the difference between a restored image and the ground truth, to evaluate the performance of the restoration algorithm, etc. Its calculation formula is:
[0057] Where MSE is the mean square error of the two images pixel by pixel.
[0058] In this embodiment, the structural similarity index SSIM is based on the perception model of the human visual system and is an indicator used to measure the similarity of two images in terms of brightness, contrast and structure. Unlike PSNR, SSIM is closer to the perception of the human visual system and can more accurately reflect the image quality. The formula of SSIM can be expressed as:
[0059] in, and are the average values of image x and y in the local window, indicating the brightness level of the image, and are the variances of images x and y in the local window, indicating the contrast of the image, is the covariance of images x and y in the local window, indicating the structural similarity of the images, , , usually K1=0.01, K2=0.03, L is the dynamic range of pixel values.
[0060] The technical effects achieved by this application are as follows: 1. The deep learning algorithm based on the multi-scale state space model is used as the core to perform image deraining. The state space model is used to model global information, and the multi-scale complementary information is effectively utilized in combination with the multi-scale framework. The cross-scale complementary features are explicitly mined to more accurately remove rain streaks, and high-quality clear images are gradually restored from rainy images, ensuring the clarity and visual quality of the derained images.
[0061] 2. By combining the multi-scale 2D selective scanning strategy with the state space model, the computational complexity is greatly reduced compared to the Transformer model method used in the existing technology. While maintaining high performance, it is more suitable for deployment on resource-constrained mobile or embedded devices.
[0062] See also Figure 5 FIG. 1 is a schematic diagram of a structure of an image rain removal system based on a multi-scale state space model provided by an embodiment of the present invention, including: An acquisition unit, used for acquiring an image with rain streaks, and decomposing the image with rain streaks into a multi-scale pyramid image by downsampling; A reconstruction unit, used for inputting the image with rain streaks and the multi-scale pyramid image into a pre-constructed multi-scale state space model, respectively, and obtaining a reconstructed residual image after performing feature extraction and image reconstruction through a convolutional layer and an encoder-decoder network, wherein the encoder-decoder network includes a multi-scale Mamba block and a frequency feature enhancement module; The obtaining unit is used to add the reconstructed residual image and the image with rain streaks to obtain a target rain-free image.
[0063] Figure 6 It is a schematic diagram of the hardware structure of an electronic device for implementing various embodiments of the present invention.
[0064] The product price version management method for public cloud scenarios provided in the embodiment of the present application can be applied to electronic devices. It can be understood by those skilled in the art that the electronic device structure involved in the embodiment of the present invention does not constitute a limitation on the electronic device, and the electronic device may include more or fewer components than shown, or combine certain components, or arrange different components. In an embodiment of the present invention, electronic devices include but are not limited to laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described and / or required herein.
[0065] The electronic device may include a processor, an external memory interface, an internal memory, a universal serial bus (USB) interface, a charging management module, a power management module, a battery, a wireless communication module, an audio module, a speaker, a microphone, a sensor module, buttons, a camera, a display, and a SIM card interface, etc.
[0066] It is to be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0067] The processor may include one or more processing units, for example, the processor may include a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated into one or more processors.
[0068] The processor can be the nerve center and command center of the electronic device. The controller can generate an operation control signal according to the instruction operation code and timing signal to complete the control of fetching and executing instructions.
[0069] A memory may also be provided in the processor for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. The memory may store instructions or data that the processor has just used or is cyclically used. If the processor needs to use the instruction or data again, it may be directly called from the memory. This avoids repeated access, reduces the waiting time of the processor, and thus improves system efficiency.
[0070] The external memory interface can be used to connect an external memory card, such as a MicroSD card, to expand the storage capacity of the electronic device. The external memory card communicates with the processor through the external memory interface to implement data storage functions. For example, files such as music and videos can be saved in the external memory card.
[0071] The internal memory can be used to store computer executable program codes, which include instructions. The processor executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory. The internal memory may include a program storage area and a data storage area. The internal memory may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0072] The wireless communication function of an electronic device can be realized through an antenna, a wireless communication module, a modem processor, and a baseband processor.
[0073] The wireless communication module can provide wireless communication solutions for electronic devices, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc.
[0074] Electronic devices can implement audio functions, etc. through audio modules, speakers, receivers, microphones, headphone jacks, and application processors.
[0075] Electronic devices can achieve shooting functions through ISP, camera, video codec, GPU, display and application processor.
[0076] Electronic devices can achieve display functions through GPU, display screen and application processor.
[0077] The GPU is a microprocessor for image processing that connects the display screen and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor may include one or more GPUs that execute program instructions to generate or change display information.
[0078] The display screen is used to display images, videos, etc. The display screen includes a display panel.
[0079] The storage medium provided in the present application stores a program product that can implement a product price version management method for a public cloud scenario.
[0080] The image deraining method based on the multi-scale state space model includes: obtaining an image with rain streaks, and decomposing the image with rain streaks into a multi-scale pyramid image by downsampling; respectively inputting the image with rain streaks and the multi-scale pyramid image into a pre-constructed multi-scale state space model, performing feature extraction and image reconstruction through a convolutional layer and an encoder-decoder network, and obtaining a reconstructed residual image, wherein the encoder-decoder network includes a multi-scale Mamba block and a frequency feature enhancement module; and adding the reconstructed residual image to the image with rain streaks to obtain a target derained image.
[0081] In some possible implementations, the subject matter of the present disclosure, namely, a product price version management method and system for public cloud scenarios, can be implemented in the form of a program product, which includes a program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps described in the above “Exemplary Method” section of this specification according to various exemplary implementations of the present disclosure.
[0082] The storage medium of the present disclosure can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0083] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An image deraining method based on a multi-scale state space model, characterized in that: include: Acquire an image with rain streaks, and decompose the image with rain streaks into a multi-scale pyramid image by downsampling; The image with rain streaks and the multi-scale pyramid image are respectively input into a pre-constructed multi-scale state space model, and after feature extraction and image reconstruction through a convolutional layer and an encoder-decoder network, a reconstructed residual image is obtained, wherein the encoder-decoder network includes a multi-scale Mamba block and a frequency feature enhancement module; The reconstructed residual image is added to the image with rain streaks to obtain a target rain-free image.
2. The image deraining method based on the multi-scale state space model according to claim 1, characterized in that: The multi-scale Mamba block is used to combine the multi-scale 2D selective scanning mechanism to apply geometric transformation operations to the multi-scale pyramid image to generate scanning sequences in different directions to represent global feature information; The frequency feature enhancement module is used to map the input features to the frequency domain through Fourier transform, separate the real part and the imaginary part and splice them into channel dimensions, extract the frequency domain features through convolution and nonlinear activation, and restore them to the spatial domain through inverse Fourier transform to realize the extraction of deep features.
3. The image deraining method based on the multi-scale state space model according to claim 2, characterized in that: The multi-scale 2D selective scanning mechanism is expressed according to the following formula: Where k represents the scale, k = 1 represents a small scale, k = 2 represents a medium scale, k = 4 represents a large scale, X represents the input image, Stack() represents the stacking operation of the image, T() represents the transposition operation of the image, F() represents the flipping operation of the pixel, and Cat() represents the splicing of the image.
4. The image deraining method based on a multi-scale state space model according to claim 1, characterized in that: The encoder-decoder network also includes a gated feature fusion module for adaptively aggregating complementary features of multi-scale branches.
5. The image deraining method based on the multi-scale state space model according to claim 4, characterized in that: The image with rain streaks and the multi-scale pyramid image are respectively input into a pre-built multi-scale state space model, and after feature extraction and image reconstruction through a convolutional layer and an encoder-decoder network, a reconstructed residual image is obtained, including: Inputting the image with rain streaks and the multi-scale pyramid image into the convolution layer of the pre-built multi-scale state space model respectively to extract shallow features; The shallow features are input into the multi-scale Mamba block in the encoder-decoder network of each scale branch, and two parallel branches are processed, and the features of the two branches are aggregated and output through the Hadamard product; wherein, in the first branch, the feature channel is processed by deep convolution and SiLU activation function, and combined with the multi-scale 2D selective scanning strategy and layer normalization operation, and in the second branch, layer normalization and SiLU activation operations are performed; The features of different scales are aggregated through the gated feature fusion module to obtain the aggregated features; The aggregated features are input into the decoder network for image reconstruction to obtain a reconstructed residual image; wherein the network architecture of the multi-scale state space model includes three scale branches, each scale branch includes an encoder and a decoder.
6. The image deraining method based on a multi-scale state space model according to claim 1, characterized in that: The features of different scales are aggregated through the gated feature fusion module, and the aggregated features are obtained according to the following formula: ; ; In the formula, Input features for different scales, express Activation function, represents 3×3 convolution, represents 1×1 convolution, ⊙ represents element-wise multiplication, represents the input features, represents the features after aggregation, Represents a gating unit.
7. The image deraining method based on a multi-scale state space model according to claim 1, characterized in that: The loss function used by the multi-scale state-space model is as follows: In the formula, represents the Charbonnier loss, represents the edge loss, represents the frequency loss, , and Represent weights respectively.
8. An image deraining system based on a multi-scale state space model, characterized in that: include: An acquisition unit, used for acquiring an image with rain streaks, and decomposing the image with rain streaks into a multi-scale pyramid image by downsampling; A reconstruction unit, used for inputting the image with rain streaks and the multi-scale pyramid image into a pre-constructed multi-scale state space model, respectively, and obtaining a reconstructed residual image after performing feature extraction and image reconstruction through a convolutional layer and an encoder-decoder network, wherein the encoder-decoder network includes a multi-scale Mamba block and a frequency feature enhancement module; The obtaining unit is used to add the reconstructed residual image and the image with rain streaks to obtain a target rain-free image.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the image deraining method based on the multi-scale state space model are implemented as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the image deraining method based on a multi-scale state space model are implemented as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Image encoding / decoding method and apparatus therefor
CN114760470A
Image blind deblurring method based on depth residual Fourier transform
CN114897741A
Multisource remote sensing image semantic segmentation method based on Transform, Mama and diffusion model
CN119152205A
Seismic image super-resolution reconstruction method based on Mama
CN119624773A
Infrared and visible light remote sensing image deep learning fusion method based on visual state space model
CN119671863A
Cited By
Remote sensing image segmentation method based on light visual scanning and frequency domain discrimination feedforward
CN121999216A
Night image rain removal method based on refined illumination modeling
CN122243805A
Vehicle-mounted visual perception enhancement method based on frequency step compensation state space model
CN122416411A
Vehicle-mounted visual perception enhancement method based on frequency step compensation state space model
CN122416411B