Low-light image restoration method based on multi-state visual angle RWKV model
Through the low-light image restoration method based on the multi-state viewing angle RWKV model, combined with brightness adaptive normalization, multi-state aggregation mechanism and state-aware selective fusion module, the problem of multiple coupling degradation in low-light environments is solved, and high-quality image restoration and dynamic adaptability are achieved.
Patent Information
- Application Number
- CN202510037210.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-06
AI Technical Summary
The existing low-light image restoration methods are difficult to effectively deal with multiple coupling degradation of images in low-light environments, resulting in significant decline in image quality and insufficient adaptability in dynamic degradation scenarios.
The low-light image restoration method based on the multi-state viewing angle RWKV model is adopted, combining brightness adaptive normalization, multi-state aggregation mechanism and state-aware selective fusion module, normalization parameters are dynamically adjusted, subtle degradation changes are captured, and multi-state features are selectively fused to reduce artifact generation.
It realizes high-quality restoration of images in low-light environments, significantly improves the retention and restoration of details, can flexibly adapt to dynamic multiple coupling degradation conditions, and reduces the amount of model parameters and computing resource requirements.
Smart Images

Figure CN119941581A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of image and video processing and computer vision technology, and in particular to a low-light image restoration method based on a multi-state perspective RWKV model. Background Art
[0002] With the rapid development of science and technology and the continuous advancement of camera technology, the application scenarios of low-light images in daily life and professional fields have increased significantly. Whether it is night photography, indoor shooting, or security monitoring and medical image capture under complex low-light conditions, the demand for image quality has shown a higher standard. However, images taken in low-light environments are usually affected by multiple degradation factors, resulting in a significant decrease in visual quality. Common degradation phenomena include increased noise, image blur, color distortion, and insufficient contrast. These types of degradation usually do not exist in isolation, but are coupled with each other, further exacerbating the degradation of image quality.
[0003] The generation of coupling degradation mainly stems from the interaction between low-light environment and multiple factors during the shooting process. For example, in order to capture more light under low-light conditions, the camera usually needs to increase the sensitivity (ISO). Although this improves the image brightness to a certain extent, it is also accompanied by the introduction of a large amount of random noise. At the same time, long exposure or device jitter will cause image blur, which superimposes noise and blur. In addition, the color information in low-light environment is weak, and the superimposed effects of noise and blur often lead to color distortion and insufficient contrast. The mutual coupling of this multiple degradation not only significantly reduces the visual quality of the image, but also has a negative impact on many practical application tasks. For example, in the field of autonomous driving, coupling degradation may lead to obstacle detection errors or road sign recognition failures, endangering driving safety; in the field of security monitoring, noise and blur cover up key details and weaken the accuracy of threat identification; in medical imaging diagnosis, the subtle features of tissue structure may not be effectively distinguished due to coupling degradation, thus affecting the accuracy of diagnosis and treatment effects.
[0004] To cope with multiple degradations in low-light environments, most existing low-light image enhancement methods optimize for a single degradation, or handle multiple degradations through a simple combination strategy of a low-light enhancement model and a deblurring model. For example, many models focus on improving brightness and suppressing noise at the same time, but this usually creates new problems: in the process of large-scale denoising, detail information is often over-smoothed, resulting in loss of image texture details. In addition, during the brightness enhancement process, high-sensitivity noise may be further enhanced due to signal amplification, making the noise more conspicuous, especially in dark areas. Deblurring techniques attempt to enhance image details by restoring motion or focal blur in the image. However, when dealing with low-light environments, blur and noise are often intertwined. Traditional deblurring methods cannot effectively separate noise while removing blur. Instead, they may introduce additional artifacts or unnatural texture enhancement, resulting in problems such as edge ghosting or discontinuous details in the image.
[0005] Solutions for joint low-light enhancement and deblurring usually adopt a strategy of processing the two tasks separately, for example, using a low-light enhancement model to increase brightness and then inputting the output image into a deblurring model. However, this pipelined processing method has significant defects: first, the models at different stages lack information sharing, resulting in the inability to fully coordinate and optimize noise and blur when they are processed separately at different stages; second, this series structure is not adaptable enough to dynamic degradation combinations. For example, in actual scenes, changes in illumination, motion, and noise are often dynamic and unpredictable. The fixed combination of task processing methods cannot flexibly cope with these complex situations, resulting in the difficulty of balancing the enhancement effect between brightness, details, and naturalness. In addition, this method may further degrade image quality due to the propagation of accumulated errors, and the output may contain residual noise or blurred areas, or introduce over-enhancement artifacts.
[0006] In summary, existing methods have significant limitations in dealing with coupled degradation, especially in real scenes, where the complex combination of degradation types such as noise, blur and color distortion is extremely unpredictable. Traditional methods are difficult to achieve ideal results in balancing multiple degradation processing, and there is an urgent need for a comprehensive solution that can flexibly adapt to multiple dynamic coupled degradations. Summary of the invention
[0007] In view of this, the purpose of the present invention is to provide a low-light image restoration method based on a multi-state perspective RWKV model. The method combines brightness adaptive normalization, dynamically adjusts the normalization parameters according to the state between stages, and realizes adaptive brightness adjustment for different lighting scenes; adopts a multi-state aggregation mechanism, aggregates multiple states within the stage through an exponential moving average strategy, captures subtle degradation changes, reduces information loss, and thus achieves a balance between detail preservation and degradation suppression; and dynamically aligns and fuses multi-state features through a state-aware selective fusion module, selectively integrates the context information of each stage of the encoder, reduces artifact generation and improves detail restoration effects.
[0008] To achieve the above object, the present invention adopts the following technical solution: a low-light image restoration method based on a multi-state perspective RWKV model, comprising the following steps:
[0009] Step A: preprocess the input image, including image pairing and data augmentation operations, to construct a training data set;
[0010] Step B, design a low-light image restoration network of a multi-state view RWKV model, which consists of an input mapping layer, a three-stage encoder consisting of several RWKV blocks, a state-aware selective fusion module, a three-stage decoder consisting of several RWKV blocks, and an output mapping layer;
[0011] Step C, designing a loss function for optimizing the restoration network described in step B;
[0012] Step D: Using the training data set constructed in step A, training a low-light image restoration network based on a multi-state perspective RWKV model;
[0013] Step E: input the image to be tested into the trained restoration network to generate a restored image under normal lighting conditions.
[0014] In a preferred embodiment, the specific implementation steps of step A are as follows:
[0015] Step A1: Pair the normal illumination image with the low light image, wherein the normal illumination image is used as the label image;
[0016] Step A2: For each training paired image, randomly select one of the following eight data augmentation methods for processing: keep the original image unchanged, flip vertically, rotate 90 degrees clockwise, rotate 90 degrees clockwise and then flip vertically, rotate 180 degrees, rotate 180 degrees and then flip vertically, rotate 270 degrees, and rotate 270 degrees and then flip vertically.
[0017] In a preferred embodiment, the specific implementation steps of step B are as follows:
[0018] Step B1: Design the input mapping layer, which consists of a convolution layer with a convolution kernel size of 3×3 and a stride of 1, to achieve low-light input images. The feature extraction of Where H and W are the height and width of the low-light image, respectively, and C represents the number of channels for extracting features;
[0019] Step B2: Design a three-stage encoder. Each stage of the encoder consists of several RWKV blocks and a downsampling layer. The number of RWKV blocks in each stage is N. 1 , and each RWKV block consists of a multi-state spatial mixing submodule and a multi-state channel mixing submodule. The downsampling layer includes a convolution layer with a kernel size of 3×3 and a stride of 1 to double the number of channels, a LeakyReLU activation function, and a bilinear downsampling operation to halve the feature size. In particular, in the third stage encoder, the number of feature channels is not doubled. After the three-stage encoder, the feature representation is
[0020] Step B3: Design a state-aware selective fusion module to improve the traditional skip connection and avoid the problem of introducing degraded information and causing artifact generation. Specifically, first aggregate the output state features from the three encoder stages and perform perceptual analysis on the state features of multiple stages. Then, concatenate the aggregated features with the input features of the encoder of the corresponding stage in the channel dimension, and then use a convolution layer with a convolution kernel size of 3×3 and a step size of 1 to fuse the concatenated features to obtain the final fused feature representation;
[0021] Step B4: Design a three-stage decoder. Each stage of the decoder consists of an upsampling layer and several RWKV blocks. The number of RWKV blocks in each stage is N. 2 The upsampling layer includes a bilinear upsampling operation to double the feature size, and a convolution layer with a kernel size of 3×3 and a stride of 1 to halve the number of channels. In particular, corresponding to the third-stage encoder, the number of feature channels in the first-stage decoder is not halved. In addition, the RWKV block in the three-stage decoder remains consistent with the structure in step B2;
[0022] Step B5: Design the output mapping layer, use a convolution layer with a convolution kernel size of 3×3 and a step size of 1, and map the feature map output in step B4 to a residual image.
[0023] Step B6: for the residual image output from step B5 With the input image Perform residual connection and generate enhanced 3-channel image by pixel-by-pixel addition, which is expressed as
[0024] In a preferred embodiment, the specific implementation steps of step B2 are as follows:
[0025] Step B21, design a multi-state spatial mixing submodule, the core components of which include a brightness adaptive normalization layer, a multi-state aggregation layer, and a bidirectional attention mechanism. First, transform the shape of C t ×H t ×W t The input features are preprocessed into a size of P by transposing and reshaping operations. t ×C t Features of X t , where t∈{1,...,N 1}(In the decoder t∈{1,...,N 2}), P t =H t ×W t Then, the feature X is transformed into t Process and obtain features Next, X s The acceptance matrix R is generated by three parallel linear projection layers s , key matrix K s Sum value matrix V s Finally, the bidirectional attention mechanism is used to calculate K s and V s The global attention weights wkv between them are activated by sigmoid and the acceptance matrix R s , to modulate the output. So, the output characteristic It can be expressed as:
[0026]
[0027] Among them, σ represents the sigmoid activation function, is the learnable parameter of the linear projection layer, wkv=BiWKV(K s ,V s ) represents the bidirectional attention proposed in Vision-RWKV (DuanY, WangW, Chen Z, et al. Vision-rwkv: Efficient and scalable visual perception with rwkv-like architectures. arXiv preprint arXiv:2403.02308, 2024.), which converts K s and Vs Calculate global attention as input;
[0028] Step B22, design a multi-state channel mixing submodule, the core components of which include a brightness adaptive normalization layer, a multi-state aggregation layer, and a channel modulation operation. First, the output features of the multi-state spatial mixing submodule are processed by the brightness adaptive normalization layer and the multi-state aggregation layer in turn to obtain a new feature representation Then, X c Generate the acceptance matrix R through two parallel linear projection layers c and key matrix K c Next, the channel modulation factor K is calculated based on Rc using the sigmoid activation function. c The mapped features are then mapped through a SquaredReLU activation function and a linear projection layer. Finally, the mapped features are multiplied by the channel modulation factor and output through a linear projection layer. The formula is as follows:
[0029]
[0030] in, and are the learnable parameters in the two linear projection layers;
[0031] Step B23: Based on the above, for the feature The output feature X' after being processed by the multi-state space mixing submodule and the multi-state channel mixing submodule t It can be expressed as:
[0032] X' t =SpatialMix(X t )+ChannelMix(X t +SpatialMix(X t )),
[0033] Among them, SpatialMix and ChannelMix represent the multi-state spatial mixing submodule and the multi-state channel mixing submodule respectively.
[0034] In a preferred embodiment, the specific implementation steps of step B21 are as follows:
[0035] Step B211: Design a brightness adaptive normalization layer, which is implemented by the brightness adaptive normalization mechanism. t∈{1,...,N 1}, P t =H t ×W t, which is converted into features by transposing and reshaping operations for First, through the global average pooling operation, And each historical stage output Here, M i represents the output of the last RWKV block in each encoder stage and decoder stage, where i∈{1,...,T-1}. Next, to maintain the semantic integrity of the luma vector, all luma vectors are expanded to the maximum number of channels C using zero padding. max The brightness vector after zero padding is recorded as where k∈{1,...,T}, the brightness vector Then they are stacked along the state dimension. Next, a multi-core aggregation layer is used to analyze the stacked brightness maps in the channel dimension. Specifically, the aggregation layer uses one-dimensional convolution kernels of size 1×T, 3×T, and 5×T to adaptively capture the local information and brightness changes of the states between different stages. Generated brightness map It is expressed as:
[0036]
[0037] in, represents several one-dimensional convolution kernels, and [·,·] represents a stacking operation;
[0038] Step B212: Output result of the convolution kernel in step B211 Concatenate in the channel dimension and map through convolution with a kernel size of 1×1 and a step size of 1 to obtain an aggregated feature map Then, a multi-layer perceptron and tanh activation function are used to predict the brightness adjustment factor Δγ t The calculation formula of the brightness adjustment factor is expressed as:
[0039] Δγ t =tanh(MLP(X agg )),
[0040] Among them, MLP represents multi-layer perceptron, and tanh represents tanh activation function.
[0041] Step B213: adjust the brightness factor Δγ t Based on , the scaling parameter γ is updated as:
[0042]
[0043] Finally, at stage T, the input feature X t The formula for processing using the brightness adaptive normalization layer can be expressed as:
[0044]
[0045] Among them, μ t and σ t Represents the input features X t The mean and standard deviation of , β is the displacement parameter;
[0046] Step B214, design a multi-state aggregation layer, aggregate the current state features with all previous state features in the same stage through the exponential moving average method, so as to achieve interaction with historical information. Specifically, for the multi-state spatial mixing submodule of the t-th RWKV block, its multi-state aggregation layer aggregates the state features processed by the brightness adaptive normalization layer and spatial state characteristics in multiple historical stages As input. Among them, Initialize to i∈{1,...,t-1}, and then updated by the output of the multi-state aggregation layer in each spatial mixing submodule. The multi-state aggregation operation is expressed as:
[0047]
[0048] Among them, α is the attenuation factor that controls the weight of the current state and the previous aggregated state, and MSA represents the multi-state aggregation operation.
[0049] In a preferred embodiment, the specific implementation steps of step B22 are as follows:
[0050] Step B221, designing a brightness adaptive normalization layer using the same structure as step B21;
[0051] Step B222: The multi-state aggregation layer structure is the same as step B21. Further, the aggregation operation here processes the state features after the brightness adaptive normalization layer is processed. Channel status characteristics in multiple historical stages For the multi-state channel mixing submodule of the t-th RWKV block, its multi-state aggregation layer converts the state features processed by the brightness adaptive normalization layer into and channel status characteristics in multiple historical stages As input. Among them, Initialize to i∈{1,...,t-1}, and then updated by the output of the multi-state aggregation layer in each channel mixing submodule. The above aggregation process formula is:
[0052]
[0053] Among them, α is the attenuation factor that controls the weight of the current state and the previous aggregated state, and MSA represents the multi-state aggregation operation.
[0054] In a preferred embodiment, the specific implementation steps of step B3 are as follows:
[0055] Step B31, design a state-aware selective fusion module (SSF). For the output of the last RWKV block of each stage encoder Among them, i∈{1,2,3}, we first perform a channel mean operation on it to compress the channel dimension, generating To reduce the semantic interference between channels. Perform an adaptive alignment operation to adjust the spatial size of each feature according to the target size of the decoder's current fusion stage. Aligned features Stacked along the channel dimension to generate a concatenated feature map Among them, H d and W d are the height and width of the decoder features respectively. This process can be expressed as:
[0056] E A =Stack(AdaAlign(Avg(E i ))),i∈{1,2,3},
[0057] Among them, Avg represents the channel mean operation, AdaAlign represents the adaptive alignment operation, and Stack is the operation of stacking the aligned features along the channel axis;
[0058] Step B32: In order to achieve more detailed processing in the fusion of local details and broad context, a perceptual block with multiple convolution kernels is used to aggregate degraded features across different receptive fields; the multiple convolution kernels include 1×1, 3×3, and 5×5; the aggregated features E agg Then it is mapped through a convolution with a convolution kernel size of 1×1 and a step size of 1, and the spatial weight W is generated by the sigmoid function. s Furthermore, taking the fusion of the encoder's third-stage features and the decoder's first-stage input features as an example, the naive skip connection operation for the RWKV block can be extended by the state-aware selective fusion module as follows:
[0059] D' 1 =([W s ⊙E 3 ,D 1 ])W p ,
[0060] Among them, W s is the spatial weight predicted by the state-aware selective fusion module, E 3is the third stage feature of the encoder, D 1 is the input feature of the first stage of the decoder, ⊙ represents element-by-element multiplication, W p denotes the learnable parameters of the mapping convolutional layer and [·,·] denotes the concatenation operation.
[0061] In a preferred embodiment, the specific implementation of step C is as follows:
[0062] Step C: Design the loss function, using L1 loss Structural loss function and VGG perceptual loss Composition, the total objective loss function of the network It is expressed as follows:
[0063]
[0064] Among them, λ 1 , 2 and λ 3 is a parameter that balances the weights of each loss, I out represents the image restored by the low-light image restoration network based on the multi-state perspective RWKV model, I gt is the image with normal illumination; μ out and μ gt is the mean of the two images, σ out and σ gt represents the variance of the two images, c 1 and c 2 are two constant values to prevent the denominator from being 0; ||·|| 2 Indicates the calculation of mean square error, VGG 3,8,15 (·) indicates that the features of layer 3, layer 8, and layer 15 are extracted using the VGG-16 classification model pre-trained on the ImageNet dataset.
[0065] In a preferred embodiment, the specific implementation steps of step D are as follows:
[0066] Step D1, randomly divide the training data set obtained in step A into several batches, each batch contains N pairs of images;
[0067] Step D2: Input the degraded low-light image I into the low-light image restoration network based on the multi-state perspective RWKV model described in step B to generate a restored image I out , and calculate the loss according to the formula in step C
[0068] Step D3, calculate the gradient of network parameters based on the loss through the back propagation method, and update the network parameters using the Adam optimization algorithm;
[0069] Step D4: Repeat steps D1 to D3 in batches until the training of the low-light image restoration network model based on the multi-state perspective RWKV model is completed.
[0070] In a preferred embodiment, the specific implementation steps of step E are as follows:
[0071] Step E1: Input the low-light image with a size of 3×H×W into the low-light image restoration network based on the multi-state perspective RWKV model described in step B to obtain a restored image I with a size of 3×H×W. out .
[0072] Compared with the prior art, the present invention has the following beneficial effects: First, by introducing a multi-state information perception mechanism, the model can effectively cope with the problem of dynamic coupling degradation in low-light environments, overcoming the inherent conflicts and limitations of existing methods in dealing with multiple coupling degradations such as noise, blur and brightness enhancement. Secondly, the brightness adaptive normalization mechanism enables the model to adapt to the lighting changes of different scenes by dynamically adjusting the normalization parameters. Thirdly, the multi-state aggregation mechanism aggregates multiple state information through an exponential moving average strategy, effectively captures subtle degradation changes, reduces information loss, and improves detail retention capabilities. Finally, the state-aware selective fusion module solves the degradation information problem that may be caused by the traditional naive jump connection method by dynamically aligning and fusing multi-state features, significantly reducing the generation of artifacts and improving the detail restoration effect. Compared with existing existing methods, the present invention significantly reduces the amount of model parameters and computing resource requirements, and provides an efficient, flexible and adaptable low-light image restoration solution for complex coupled degradation scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] Figure 1 It is a flow chart of the implementation of the method in the preferred embodiment of the present invention.
[0074] Figure 2 It is a structural diagram of a low-light image restoration network based on a multi-state perspective RWKV model in a preferred embodiment of the present invention.
[0075] Figure 3 It is a schematic diagram of a brightness adaptive normalization mechanism in a preferred embodiment of the present invention.
[0076] Figure 4 It is a structural diagram of the state-aware selective fusion module in the preferred embodiment of the present invention. DETAILED DESCRIPTION
[0077] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0078] It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.
[0079] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or their combinations.
[0080] The present invention provides a low-light image restoration method based on a multi-state viewing angle RWKV model. Figure 1-4 As shown, the following steps are included:
[0081] Step A: preprocess the input image, including image pairing and data augmentation operations, to construct a training data set;
[0082] Step B, design a low-light image restoration network of a multi-state view RWKV model, which consists of an input mapping layer, a three-stage encoder consisting of several RWKV blocks, a state-aware selective fusion module, a three-stage decoder consisting of several RWKV blocks, and an output mapping layer;
[0083] Step C, designing a loss function for optimizing the restoration network described in step B;
[0084] Step D: Using the training data set constructed in step A, training a low-light image restoration network based on a multi-state perspective RWKV model;
[0085] Step E: input the image to be tested into the trained restoration network to generate a restored image under normal lighting conditions.
[0086] Furthermore, the step A comprises the following steps:
[0087] Step A1: Pair the normal illumination image with the low light image, wherein the normal illumination image is used as the label image;
[0088] Step A2: For each training paired image, randomly select one of the following eight data augmentation methods for processing: keep the original image unchanged, flip vertically, rotate 90 degrees clockwise, rotate 90 degrees clockwise and then flip vertically, rotate 180 degrees, rotate 180 degrees and then flip vertically, rotate 270 degrees, and rotate 270 degrees and then flip vertically.
[0089] Furthermore, the step B comprises the following steps:
[0090] Step B1: Design the input mapping layer, which consists of a convolution layer with a convolution kernel size of 3×3 and a stride of 1, to achieve low-light input images. The feature extraction of Where H and W are the height and width of the low-light image, respectively, and C represents the number of channels for extracting features;
[0091] Step B2: Design a three-stage encoder. Each stage of the encoder consists of several RWKV blocks and a downsampling layer. The number of RWKV blocks in each stage is N. 1 , and each RWKV block consists of a multi-state spatial mixing submodule and a multi-state channel mixing submodule. The downsampling layer includes a convolution layer with a kernel size of 3×3 and a stride of 1 to double the number of channels, a LeakyReLU activation function, and a bilinear downsampling operation to halve the feature size. In particular, in the third stage encoder, the number of feature channels is not doubled. After the three-stage encoder, the feature representation is
[0092] Step B3: Design a state-aware selective fusion module to improve the traditional skip connection and avoid the problem of introducing degraded information and causing artifact generation. Specifically, first aggregate the output state features from the three encoder stages and perform perceptual analysis on the state features of multiple stages. Then, concatenate the aggregated features with the input features of the encoder of the corresponding stage in the channel dimension, and then use a convolution layer with a convolution kernel size of 3×3 and a step size of 1 to fuse the concatenated features to obtain the final fused feature representation;
[0093] Step B4: Design a three-stage decoder. Each stage of the decoder consists of an upsampling layer and several RWKV blocks. The number of RWKV blocks in each stage is N. 2 The upsampling layer includes a bilinear upsampling operation to double the feature size, and a convolution layer with a kernel size of 3×3 and a stride of 1 to halve the number of channels. In particular, corresponding to the third-stage encoder, the number of feature channels in the first-stage decoder is not halved. In addition, the RWKV block in the three-stage decoder remains consistent with the structure in step B2;
[0094] Step B5: Design the output mapping layer, use a convolution layer with a convolution kernel size of 3×3 and a step size of 1, and map the feature map output in step B4 to a residual image.
[0095] Step B6: for the residual image output from step B5 With the input image Perform residual connection and generate enhanced 3-channel image by pixel-by-pixel addition, which is expressed as
[0096] Furthermore, the step B2 comprises the following steps:
[0097] Step B21, design a multi-state spatial mixing submodule, the core components of which include a brightness adaptive normalization layer, a multi-state aggregation layer, and a bidirectional attention mechanism. First, transform the shape of C t ×H t ×W t The input features are preprocessed into a size of P by transposing and reshaping operations. t ×C t Features of X t , where t∈{1,...,N 1}(In the decoder t∈{1,...,N 2}), P t =H t ×W t Then, the feature X is transformed into t Process and obtain features Next, X s The acceptance matrix R is generated by three parallel linear projection layers s , key matrix K s Sum value matrix V s Finally, the bidirectional attention mechanism is used to calculate K s and V s The global attention weights wkv between them are activated by sigmoid and the acceptance matrix R s , to modulate the output. So, the output characteristic It can be expressed as:
[0098]
[0099] Among them, σ represents the sigmoid activation function, is the learnable parameter of the linear projection layer, wkv=BiWKV(K s ,V s ) represents the bidirectional attention proposed in Vision-RWKV (DuanY, Wang W, Chen Z, et al. Vision-rwkv: Efficient and scalable visual perception with rwkv-like architectures. arXiv preprint arXiv:2403.02308, 2024.), which converts K s and Vs Calculate global attention as input;
[0100] Step B22, design a multi-state channel mixing submodule, the core components of which include a brightness adaptive normalization layer, a multi-state aggregation layer, and a channel modulation operation. First, the output features of the multi-state spatial mixing submodule are processed by the brightness adaptive normalization layer and the multi-state aggregation layer in turn to obtain a new feature representation Then, X c Generate the acceptance matrix R through two parallel linear projection layers c and key matrix K c Next, we use the sigmoid activation function in R c Based on this, the channel modulation factor, K c The mapped features are then mapped through a SquaredReLU activation function and a linear projection layer. Finally, the mapped features are multiplied by the channel modulation factor and output through a linear projection layer. The formula is as follows:
[0101]
[0102] in, and are the learnable parameters in the two linear projection layers;
[0103] Step B23: Based on the above, for the feature The output feature X' after being processed by the multi-state space mixing submodule and the multi-state channel mixing submodule t It can be expressed as:
[0104] X' t =SpatialMix(X t )+ChannelMix(X t +SpatialMix(X t )),
[0105] Among them, SpatialMix and ChannelMix represent the multi-state spatial mixing submodule and the multi-state channel mixing submodule respectively.
[0106] Furthermore, the step B21 includes the following steps:
[0107] Step B211: Design a brightness adaptive normalization layer, which is implemented by the brightness adaptive normalization mechanism. t∈{1,...,N 1}, P t =H t ×W t, which is converted into features by transposing and reshaping operations for First, through the global average pooling operation, And each historical stage output Here, M i represents the output of the last RWKV block in each encoder stage and decoder stage, where i∈{1,...,T-1}. Next, to maintain the semantic integrity of the luma vector, all luma vectors are expanded to the maximum number of channels C using zero padding. max The brightness vector after zero padding is recorded as where k∈{1,...,T}, the brightness vector Then they are stacked along the state dimension. Next, a multi-core aggregation layer is used to analyze the stacked brightness maps in the channel dimension. Specifically, the aggregation layer uses one-dimensional convolution kernels of size 1×T, 3×T, and 5×T to adaptively capture the local information and brightness changes of the states between different stages. Generated brightness map It is expressed as:
[0108]
[0109] in, represents several one-dimensional convolution kernels, and [·,·] represents a stacking operation;
[0110] Step B212: Output result of the convolution kernel in step B211 Concatenate in the channel dimension and map through convolution with a kernel size of 1×1 and a step size of 1 to obtain an aggregated feature map Then, a multi-layer perceptron and tanh activation function are used to predict the brightness adjustment factor Δγ t The calculation formula of the brightness adjustment factor is expressed as:
[0111] Δγ t =tanh(MLP(X agg )),
[0112] Among them, MLP represents multi-layer perceptron, and tanh represents tanh activation function.
[0113] Step B213: adjust the brightness factor Δγ t Based on , the scaling parameter γ is updated as:
[0114]
[0115] Finally, at stage T, the input feature X t The formula for processing using the brightness adaptive normalization layer can be expressed as:
[0116]
[0117] Among them, μ t and σ t Represents the input features X t The mean and standard deviation of , β is the displacement parameter;
[0118] Step B214, design a multi-state aggregation layer, aggregate the current state features with all previous state features in the same stage through the exponential moving average method, so as to achieve interaction with historical information. Specifically, for the multi-state spatial mixing submodule of the t-th RWKV block, its multi-state aggregation layer aggregates the state features processed by the brightness adaptive normalization layer and spatial state characteristics in multiple historical stages As input. Among them, Initialize to i∈{1,...,t-1}, and then updated by the output of the multi-state aggregation layer in each spatial mixing submodule. The multi-state aggregation operation is expressed as:
[0119]
[0120] Among them, α is the attenuation factor that controls the weight of the current state and the previous aggregated state, and MSA represents the multi-state aggregation operation.
[0121] Furthermore, the step B22 includes the following steps:
[0122] Step B221, designing a brightness adaptive normalization layer using the same structure as step B21;
[0123] Step B222: The multi-state aggregation layer structure is the same as step B21. Further, the aggregation operation here processes the state features after the brightness adaptive normalization layer is processed. Channel status characteristics in multiple historical stages For the multi-state channel mixing submodule of the t-th RWKV block, its multi-state aggregation layer converts the state features processed by the brightness adaptive normalization layer into and channel status characteristics in multiple historical stages As input. Among them, Initialize to i∈{1,...,t-1}, and then updated by the output of the multi-state aggregation layer in each channel mixing submodule. The above aggregation process formula is:
[0124]
[0125] Among them, α is the attenuation factor that controls the weight of the current state and the previous aggregated state, and MSA represents the multi-state aggregation operation.
[0126] Further, step B3 is implemented as follows:
[0127] Step B31, design a state-aware selective fusion module (SSF). For the output of the last RWKV block of each stage encoder Among them, i∈{1,2,3}, we first perform a channel mean operation on it to compress the channel dimension, generating To reduce the semantic interference between channels. Perform an adaptive alignment operation to adjust the spatial size of each feature according to the target size of the decoder's current fusion stage. Aligned features Stacked along the channel dimension to generate a concatenated feature map Among them, H d and W d are the height and width of the decoder features respectively. This process can be expressed as:
[0128] E A =Stack(AdaAlign(Avg(E i ))),i∈{1,2,3},
[0129] Among them, Avg represents the channel mean operation, AdaAlign represents the adaptive alignment operation, and Stack is the operation of stacking the aligned features along the channel axis;
[0130] Step B32: In order to achieve more detailed processing in the fusion of local details and broad context, a perceptual block with multiple convolution kernels is used to aggregate degraded features across different receptive fields; the multiple convolution kernels include 1×1, 3×3, and 5×5; the aggregated features E agg Then it is mapped through a convolution with a convolution kernel size of 1×1 and a step size of 1, and the spatial weight W is generated by the sigmoid function. s Furthermore, taking the fusion of the encoder's third-stage features and the decoder's first-stage input features as an example, the naive skip connection operation for the RWKV block can be extended by the state-aware selective fusion module as follows:
[0131] D' 1 =([W s ⊙E 3 ,D 1 ])W p ,
[0132] Among them, W s is the spatial weight predicted by the state-aware selective fusion module, E 3 is the third stage feature of the encoder, D1 is the input feature of the first stage of the decoder, ⊙ represents element-by-element multiplication, W p denotes the learnable parameters of the mapping convolutional layer and [·,·] denotes the concatenation operation.
[0133] Further, step C is implemented as follows:
[0134] Step C: Design the loss function, using L1 loss Structural loss function and VGG perceptual loss Composition, the total objective loss function of the network It is expressed as follows:
[0135]
[0136] Among them, λ 1 , 2 and λ 3 is a parameter that balances the weights of each loss, I out represents the image restored by the low-light image restoration network based on the multi-state perspective RWKV model, I gt is the image with normal illumination; μ out and μ gt is the mean of the two images, σ out and σ gt represents the variance of the two images, c 1 and c 2 are two constant values to prevent the denominator from being 0; ||·|| 2 Indicates the calculation of mean square error, VGG 3,8,15 (·) indicates that the features of layer 3, layer 8, and layer 15 are extracted using the VGG-16 classification model pre-trained on the ImageNet dataset.
[0137] Further, the step D is implemented as follows:
[0138] Step D1, randomly divide the training data set obtained in step A into several batches, each batch contains N pairs of images;
[0139] Step D2: Input the degraded low-light image I into the low-light image restoration network based on the multi-state perspective RWKV model described in step B to generate a restored image I out , and calculate the loss according to the formula in step C
[0140] Step D3, calculate the gradient of network parameters based on the loss through the back propagation method, and update the network parameters using the Adam optimization algorithm;
[0141] Step D4: Repeat steps D1 to D3 in batches until the training of the low-light image restoration network model based on the multi-state perspective RWKV model is completed.
[0142] Further, the step E is implemented as follows:
[0143] Step E1: Input the low-light image with a size of 3×H×W into the low-light image restoration network based on the multi-state perspective RWKV model described in step B to obtain a restored image I with a size of 3×H×W. out .
[0144] The above are preferred embodiments of the present invention. Any changes made according to the technical solution of the present invention, as long as the resulting functions do not exceed the scope of the technical solution of the present invention, belong to the protection scope of the present invention.
[0145] The present invention aims to solve the problem that the existing low-light image restoration methods are only designed for preset degradation conditions and are difficult to achieve high-quality restoration in dynamic coupled degradation scenarios. To this end, a low-light image restoration method based on a multi-state perspective RWKV model is designed. First, a brightness adaptive normalization mechanism is designed. This mechanism draws on the principle of adaptive adjustment of the human eye pupil. For images under different lighting conditions, the normalization parameters are dynamically adjusted to achieve adaptive brightness adjustment. Secondly, by designing a multi-state aggregation mechanism, the exponential moving average strategy is used to aggregate state information in multiple stages, effectively capturing subtle degradation changes, reducing information loss, and thus improving detail retention and restoration effects. Finally, a state-aware selective fusion module is designed. This module solves the problem that traditional jump connections may introduce degradation information by dynamically aligning and fusing multi-state features, significantly reducing artifact generation, and further improving detail restoration capabilities. In summary, the multi-state perspective RWKV model proposed in the present invention effectively overcomes the limitations of existing methods, provides an efficient and flexible low-light image restoration solution, and can significantly improve image quality under dynamic multi-coupled degradation conditions.
Claims
1. A low-light image restoration method based on a multi-state perspective RWKV model, characterized in that: The steps include: Step A: preprocess the input image, including image pairing and data augmentation operations, to construct a training data set; Step B, design a low-light image restoration network of a multi-state view RWKV model, including an input mapping layer, a three-stage encoder consisting of several RWKV blocks, a state-aware selective fusion module, a three-stage decoder consisting of several RWKV blocks, and an output mapping layer; Step C, designing a loss function for optimizing the restoration network described in step B; Step D: Using the training data set constructed in step A, training a low-light image restoration network based on a multi-state perspective RWKV model; Step E: input the image to be tested into the trained restoration network to generate a restored image under normal lighting conditions.
2. The low-light image restoration method based on the multi-state perspective RWKV model according to claim 1, characterized in that: The specific implementation steps of step A are as follows: Step A1: Pair the normal illumination image with the low light image, wherein the normal illumination image is used as the label image; Step A2: For each training paired image, randomly select one of the following eight data augmentation methods for processing: keep the original image unchanged, flip vertically, rotate 90 degrees clockwise, rotate 90 degrees clockwise and then flip vertically, rotate 180 degrees, rotate 180 degrees and then flip vertically, rotate 270 degrees, and rotate 270 degrees and then flip vertically.
3. The low-light image restoration method based on the multi-state perspective RWKV model according to claim 1, characterized in that: The specific implementation steps of step B are as follows: Step B1: Design the input mapping layer, which consists of a convolution layer with a convolution kernel size of 3×3 and a stride of 1, to achieve low-light input images. The feature extraction of Where H and W are the height and width of the low-light image, respectively, and C represents the number of channels for extracting features; Step B2, design a three-stage encoder, each stage of the encoder consists of several RWKV blocks and a downsampling layer, where the number of RWKV blocks contained in each stage is N1, and each RWKV block consists of a multi-state spatial mixing submodule and a multi-state channel mixing submodule; the downsampling layer includes a convolution layer with a convolution kernel size of 3×3 and a stride of 1, which is used to double the number of channels, a LeakyReLU activation function, and a bilinear downsampling operation to halve the feature size; In the third stage encoder, the number of feature channels is not doubled; after the three-stage encoder processing, the feature representation is Step B3, design a state-aware selective fusion module to improve the traditional skip connection; first, aggregate the output state features from the three encoder stages, and perform perceptual analysis on the state features of multiple stages; then, concatenate the aggregated features with the input features of the encoder of the corresponding stage in the channel dimension, and then use a convolution layer with a convolution kernel size of 3×3 and a stride of 1 to fuse the concatenated features to obtain the final fused feature representation; Step B4, design a three-stage decoder, where each stage of the decoder consists of an upsampling layer and several RWKV blocks, where the number of RWKV blocks in each stage is N2; the upsampling layer includes a bilinear upsampling operation to double the feature size, and a convolution layer with a convolution kernel size of 3×3 and a stride of 1 to halve the number of channels; corresponding to the third-stage encoder, the number of feature channels of the first-stage decoder is not halved; in addition, the RWKV block in the three-stage decoder is consistent with the structure in step B2; Step B5: Design the output mapping layer, use a convolution layer with a convolution kernel size of 3×3 and a step size of 1, and map the feature map output in step B4 to a residual image. Step B6: for the residual image output from step B5 With the input image Perform residual connection and generate enhanced 3-channel image by pixel-by-pixel addition, which is expressed as 4. The low-light image restoration method based on the multi-state perspective RWKV model according to claim 3, characterized in that: The specific implementation steps of step B2 are as follows: Step B21, design a multi-state spatial mixing submodule. The core components of the multi-state spatial mixing submodule include a brightness adaptive normalization layer, a multi-state aggregation layer, and a bidirectional attention mechanism. First, convert the shape of C t ×H t ×W t The input features are preprocessed into a size of P by transposing and reshaping operations. t ×C t Features of X t , where t∈{1,...,N1}, where in the decoder t∈{1,...,N2}, P t =H t ×W t ; Then, the feature X is processed by the adaptive normalization layer and the multi-state aggregation layer in turn. t Process and obtain features Next, X s The acceptance matrix R is generated by three parallel linear projection layers s , key matrix K s Sum value matrix V s ; Finally, the bidirectional attention mechanism is used to calculate K s and V s The global attention weights wkv between them are activated by sigmoid and the acceptance matrix R s , modulate the output; Output Features The formula is: Among them, σ represents the sigmoid activation function, is the learnable parameter of the linear projection layer, wdv=BiWKV(K s ,V s ) represents the bidirectional attention proposed in Vision-RWKV, K s and V s Calculate global attention as input; Step B22, design a multi-state channel mixing submodule. The core components of the multi-state channel mixing submodule include a brightness adaptive normalization layer, a multi-state aggregation layer, and a channel modulation operation. First, the output features of the multi-state spatial mixing submodule are processed by the brightness adaptive normalization layer and the multi-state aggregation layer in turn to obtain a new feature representation. Then, X c Generate the acceptance matrix R through two parallel linear projection layers c and key matrix K c ; Next, use the sigmoid activation function in R c Based on this, the channel modulation factor, K c The mapped features are then mapped through the SquaredReLU activation function and a linear projection layer; finally, the mapped features are multiplied by the channel modulation factor and the final features are output through a linear projection layer. The formula is as follows: in, and are the learnable parameters in the two linear projection layers; Step B23: For features The output feature X' after being processed by the multi-state space mixing submodule and the multi-state channel mixing submodule t It is expressed as: X' t =SpatialMix(X t )+ChannelMix(X t +SpatialMix(X t )), Among them, SpatialMix and ChannelMix represent the multi-state spatial mixing submodule and the multi-state channel mixing submodule respectively.
5. The low-light image restoration method based on the multi-state perspective RWKV model according to claim 4, characterized in that: The specific implementation steps of step B21 are as follows: Step B211, design a brightness adaptive normalization layer, which is implemented by the brightness adaptive normalization mechanism; for the input of the current stage T P t =H t ×W t , through the transpose and reshape operations Convert to Features for First, through the global average pooling operation, And each historical stage output Extract the brightness vector representation from M; i represents the output of the last RWKV block in each encoder stage and decoder stage, i∈{1,...,T-1}; then, to maintain the semantic integrity of the brightness vector, all brightness vectors are expanded to the maximum number of channels C using zero padding operations. max The brightness vector after zero filling is recorded as where k∈{1,...,T}, the brightness vector Then, the images are stacked along the state dimension. Next, the stacked brightness images are analyzed in the channel dimension using a multi-core aggregation layer. The aggregation layer uses one-dimensional convolution kernels of size 1×T, 3×T, and 5×T to adaptively capture the local information and brightness changes of the states between different stages. The generated brightness images It is expressed as: in, represents several one-dimensional convolution kernels, and [·,·] represents a stacking operation; Step B212: Output result of the convolution kernel in step B211 Concatenate in the channel dimension and map through convolution with a kernel size of 1×1 and a step size of 1 to obtain an aggregated feature map Then, a multi-layer perceptron and tanh activation function are used to predict the brightness adjustment factor Δγ t ; The calculation formula of brightness adjustment factor is expressed as: Δγ t =tanh(MLP(X agg )), Among them, MLP represents multi-layer perceptron, tanh represents tanh activation function; Step B213: adjust the brightness factor Δγ t Based on , the scaling parameter γ is updated as: Finally, at stage T, the input feature X t The formula for processing using the brightness adaptive normalization layer is expressed as: Among them, μ t and σ t Represents the input features X t The mean and standard deviation of , β is the displacement parameter; Step B214, design a multi-state aggregation layer, aggregate the current state features with all previous state features in the same stage through the exponential moving average method, and realize the interaction with historical information; for the multi-state spatial mixing submodule of the t-th RWKV block, its multi-state aggregation layer converts the state features processed by the brightness adaptive normalization layer into and spatial state characteristics in multiple historical stages As input; where Initialize to It is then updated by the output of the multi-state aggregation layer in each spatial mixing submodule; the multi-state aggregation operation is expressed as: Among them, α is the attenuation factor that controls the weight of the current state and the previous aggregated state, and MSA represents the multi-state aggregation operation.
6. The low-light image restoration method based on the multi-state perspective RWKV model according to claim 5, characterized in that: The specific implementation steps of step B22 are as follows: Step B221, designing a brightness adaptive normalization layer using the same structure as step B21; Step B222: The multi-state aggregation layer structure is the same as step B21. Furthermore, the aggregation operation processes the state features after the brightness adaptive normalization layer is processed. Channel status characteristics in multiple historical stages For the multi-state channel mixing submodule of the t-th RWKV block, its multi-state aggregation layer converts the state features processed by the brightness adaptive normalization layer into and channel status characteristics in multiple historical stages As input; where Initialize to It is then updated by the output of the multi-state aggregation layer in each channel mixing submodule; the above aggregation process formula is: Among them, α is the attenuation factor that controls the weight of the current state and the previous aggregated state, and MSA represents the multi-state aggregation operation.
7. The low-light image restoration method based on the multi-state perspective RWKV model according to claim 3, characterized in that: The specific implementation steps of step B3 are as follows: Step B31, design a state-aware selective fusion module SSF; for the output of the last RWKV block of each stage encoder Where i∈{1,2,3}, first Perform channel mean operation to compress channel dimension and generate To reduce the semantic interference between channels; then, Perform adaptive alignment to adjust the spatial size of each feature according to the target size of the decoder's current fusion stage; the aligned features Stacked along the channel dimension to generate a concatenated feature map Among them, H d and W d are the height and width of the decoder features respectively; this process is expressed by the formula: E A =Stack(AdaAlign(Avg(E i ))),i∈{1,2,3}, Among them, Avg represents the channel mean operation, AdaAlign represents the adaptive alignment operation, and Stack is the operation of stacking the aligned features along the channel axis; Step B32: In order to achieve more detailed processing in the fusion of local details and broad context, a perceptual block with multiple convolution kernels is used to aggregate degraded features across different receptive fields; the multiple convolution kernels include 1×1, 3×3, and 5×5; the aggregated features E agg Then it is mapped through a convolution with a kernel size of 1×1 and a step size of 1, and the spatial weight W is generated by the sigmoid function. s ; Furthermore, the naive skip connection operation for the RWKV block is extended by the state-aware selective fusion module as follows: D'1=([W s ⊙E3,D1])W p , Among them, W s is the spatial weight predicted by the state-aware selective fusion module, E3 is the third-stage feature of the encoder, D1 is the input feature of the first stage of the decoder, ⊙ represents element-by-element multiplication, and W p denotes the learnable parameters of the mapping convolutional layer and [·,·] denotes the concatenation operation.
8. The low-light image restoration method based on the multi-state perspective RWKV model according to claim 1, characterized in that: The specific implementation of step C is: Step C: Design the loss function, using L1 loss Structural loss function l structure and VGG perceptual loss l perceptual Composition, the total objective loss function of the network l total It is expressed as follows: l perceptual =||(VGG 3,8,15 (I out )-VGG 3,8,15 (I gt ))|| 2 Among them, λ1, λ2 and λ3 are parameters that balance the weights of each loss, I out represents the image restored by the low-light image restoration network based on the multi-state perspective RWKV model, I gt is the image with normal illumination; μ out and μ gt is the mean of the two images, σ out and σ gt represents the variance of two images, c1 and c2 are two constant values to prevent the denominator from being 0; ||·|| 2 Indicates the calculation of mean square error, VGG 3,8,15 (·) indicates that the features of layer 3, layer 8, and layer 15 are extracted using the VGG-16 classification model pre-trained on the ImageNet dataset.
9. The low-light image restoration method based on the multi-state perspective RWKV model according to claim 1, characterized in that: The specific implementation steps of step D are as follows: Step D1, randomly divide the training data set obtained in step A into several batches, each batch containing N pairs of images; Step D2: Input the degraded low-light image I into the low-light image restoration network based on the multi-state perspective RWKV model described in step B to generate a restored image I out , and calculate the loss l according to the formula in step C total ; Step D3, calculate the gradient of network parameters based on the loss through the back propagation method, and update the network parameters using the Adam optimization algorithm; Step D4: Repeat steps D1 to D3 in batches until the training of the low-light image restoration network model based on the multi-state perspective RWKV model is completed.
10. The low-light image restoration method based on the multi-state perspective RWKV model according to claim 1, characterized in that: The specific implementation steps of step E are as follows: Step E1: Input the low-light image with a size of 3×H×W into the low-light image restoration network based on the multi-state perspective RWKV model described in step B to obtain a restored image I with a size of 3×H×W. out .
Citation Information
Cited By
End-to-end automatic driving method based on linear time complexity
CN120763872A
Infrared image and visible light image fusion method and system based on RWKV feature fusion
CN120876260A
Three-dimensional medical image segmentation system and method based on three-dimensional structure enhancement
CN121121130A