Low-light raw image enhancement method based on rwkv

By using a CRF module and attention mechanism based on RWKV, the noise and dynamic range problems of low-light RAW images are solved, achieving efficient image enhancement effects suitable for image processing under low-light conditions.

CN120410900BActive Publication Date: 2025-11-07ANHUI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510533374.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-26
Publication Date
2025-11-07
Estimated Expiration
2045-04-26

AI Technical Summary

Technical Problem

In low-light conditions, RAW images captured by camera sensors are noisy and have limited dynamic range, resulting in blurred details and color distortion. Traditional methods struggle to simultaneously address brightness unevenness and noise suppression, while existing CNN-based methods consume high computational resources and are unsuitable for real-time applications.

Method used

A low-light RAW image enhancement method based on RWKV is adopted. By constructing a CRF module and combining convolutional and RWKV branches, and utilizing an omnidirectional token shift layer and Re-WKV attention mechanism, local and global features of the image are captured to achieve nonlinear enhancement.

Benefits of technology

It effectively restores local details and overall structure of low-light RAW images, improves image clarity and quality, reduces computational complexity, and is suitable for real-time applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120410900B_ABST
    Figure CN120410900B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of computer vision and image processing, and proposes a low-light RAW image enhancement method based on RWKV. This method directly uses RAW image data for low-light image enhancement, improving the performance deficiency of traditional methods in processing low-light RAW images, greatly improving the quality of low-light RAW images. The structure of the RawRWKV model in the present application is based on the U-Net architecture and introduces a CRF module, which contains a parallel convolution branch and a RWKV branch, combining local attention and global attention mechanisms, enabling it to effectively capture structural relationships in images. Meanwhile, the RWKV attention mechanism is introduced, reducing the complexity to O, greatly reducing the computational overhead. This technology has shown significant technical advantages in low-light image enhancement, with its efficient computing power, excellent image quality improvement and wide applicability, making it have important application potential in multiple fields.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision and image processing, in particular to a low-light RAW image enhancement method based on RWKV. BACKGROUND

[0002] With the rapid development of computer vision technology, digital image processing systems have been widely integrated into all aspects of daily life. However, in low-light environments, images often have problems such as insufficient brightness, color distortion, and noise interference, which directly affect the performance of downstream tasks such as target detection and security monitoring. Traditional physical methods (such as increasing the aperture and prolonging the exposure) can improve the amount of light, but are limited by device size, motion blur, and noise amplification, making it difficult to be widely applied. Therefore, algorithm enhancement has become the core direction to solve the problem of low light.

[0003] In order to enhance low-light images, traditional image processing methods mainly rely on hand-designed algorithms such as histogram equalization and Retinex theory. Although these methods can improve the contrast and brightness of images to some extent, they often cannot effectively handle complex low-light images, especially for noise and color distortion. In addition, these methods usually require manual parameter adjustment and lack adaptability. With the development of deep learning technology, methods based on convolutional neural networks (CNN) have gradually become the mainstream of low-light image enhancement. These methods learn the mapping relationship from low-light images to high-quality images, and can automatically extract image features and enhance them. However, existing CNN-based methods mainly focus on local features and are difficult to capture long-range dependencies in images, and the model is usually large, consuming high computational resources, which is not suitable for real-time applications. The Transformer can capture long-range dependencies in images through self-attention mechanisms, but its computational complexity is proportional to the square of the image resolution, resulting in low efficiency on high-resolution images.

[0004] RWKV model derived from natural language processing (NLP) is becoming an alternative to Transformer. RWKV introduces two important innovations. On the one hand, it introduces a RWKV attention mechanism that can build long-range dependencies with linear computational complexity, solving the quadratic computational complexity problem in the self-attention mechanism of Transformer. On the other hand, it introduces a token shift layer to enhance the capture of local dependencies, which is usually ignored by the standard Transformer. When existing models are migrated from natural language processing to visual tasks, they show better performance compared to visual Transformers with reduced computational complexity. In order to better utilize the spatial information in two-dimensional images, some studies have proposed a bidirectional WKV attention mechanism to capture global dependencies, and a four-way token shift mechanism to capture local context information from four different directions; based on this mechanism, a number of RWKV-based models have been developed for various visual-related tasks, including Diffusion-RWKV for image generation, RWKV-SAM for segmenting arbitrary objects, and RWKV-CLIP for visual-language representation learning. However, few studies have applied RWKV to low-level visual tasks such as low-light image enhancement. In view of this, we propose a low-light RAW image enhancement method based on RWKV. SUMMARY

[0005] The purpose of the present application is to solve the problem that under low light conditions, the RAW image collected by the camera sensor has noise and the dynamic range is limited, resulting in blurred details and color distortion; low-light image enhancement needs to handle brightness unevenness (dark part brightening, bright part detail preservation) and noise suppression simultaneously, and traditional linear transformation (such as global histogram equalization) cannot meet the non-linear requirements.

[0006] To achieve the above purpose, the present application provides a low-light RAW image enhancement method based on RWKV, comprising the following steps:

[0007] S1, adjust the size and state of the RAW image using data augmentation techniques to expand the training data set; use the original data as the original data input for the test set; sample the input image in the test set and convert it to a multi-channel input;

[0008] S2, extract local features of the multi-channel input through a convolutional embedding layer, and use an activation function to increase the nonlinearity of the model;

[0009] S3, construct a CRF module, including a parallel convolution branch and a RWKV branch, integrate the results of the two branches, and output the feature map obtained through two convolution layers and an activation layer;

[0010] S4, construct a multi-level encoder-decoder based on CRF to generate the corresponding feature map;

[0011] S5. Construct the output module, apply convolutional layers and shuffling operations to output an enhanced RGB image.

[0012] As a further improvement to this technical solution, the expression for image sampling in S1 is as follows: Perceiving low-light RAW images RAW images Corresponding size , and RAW images Height and width; set The output channel image has a size of [size missing]. The calculation formula is: ;in To output image coordinates, For input channel index; Represents RAW image In the middle, the coordinates are The pixel value with channel index 0.

[0013] As a further improvement to this technical solution, the convolutional embedding layer in S2 is implemented through... convolution kernel Slide across the channel image to cover the channel image. The local region is then processed, and each element in the convolution kernel is multiplied by the corresponding image pixel value. The products are then summed to obtain the data. This process is repeated at each position in the channel image to obtain local features.

[0014] As a further improvement to this technical solution, the expression for model nonlinearity added in S2 is as follows: ;

[0015] in It is a set function; specifically, it is; hour, The function is linearly activated with a normal slope, preserving strong signal characteristics; hour, , The slope is positive, and the slope is... Activation, applying a slope to dark features enlarge.

[0016] As a further improvement to this technical solution, the convolution branch in S3 is used to obtain local features of the image, and the corresponding expression is: The RWKV branch structure consists of two parts: Spatial Mix and Channel Mix.

[0017] The Spatial Mix component employs Layer Norm to normalize local features. Specifically, the mean of the local features is normalized to 0, and the variance is normalized to 1. The corresponding expression is: The perceptual input feature matrix is... ,in For the sample size, The feature dimension is calculated using the following formula: Normalization calculations are performed; where local features are in dimension mean Local features in dimension The variance is ; It is a constant; and These are learnable parameters; Characteristic matrix The Middle Line number The elements of the column.

[0018] As a further improvement to this technical solution, the Spatial Mix uses an omnidirectional token shift layer to adjust spatial position information. The omnidirectional token shift layer shifts and merges tokens from all directions; different branches are responsible for token shifting in different context ranges.

[0019] The omnidirectional token shift layer combines features by incorporating convolution operations of different sizes and the input itself, and its corresponding expression is:

[0020] ;

[0021] in This is the output of Layer Norm, i.e., the local features after normalization. These are weighting coefficients used to control the contribution ratio of convolutional operations at different scales and the input itself in the omnidirectional token shift layer. They are respectively represented as The convolution operation.

[0022] The beneficial effect of adopting the above-mentioned further scheme is that the omnidirectional token shifting mechanism uses convolution to move and merge adjacent tokens from all directions, realizing an accurate and efficient token shifting mechanism to aggregate local context, overcoming the problem that it can only shift tokens from a limited number of directions and cannot make full use of the inherent spatial relationships in two-dimensional images.

[0023] Based on the above technical solution, the present invention can be further improved as follows.

[0024] As a further improvement to this technical solution, the omnidirectional token shifting layer generation... The three branches extract and adjust spatial location information at different scales, among which... ;

[0025] in There are three linear projection layers, and the linear projection layers are used to... Perform a linear transformation to obtain the following results: Three branches;

[0026] After the omnidirectional token shift layer operation, and The input is fed into Re-WKV, and the result is obtained through the Re-WKV attention mechanism. The output, corresponding to the expression, is: ; ;

[0027] in This is data that integrates information from all local image feature units. Represented as The Middle Attention weight values ​​corresponding to each local image feature unit This represents the total number of feature units in a local image. For traversing 1 to Index variables for local image feature unit information; for The Middle The key value corresponding to each local image feature unit. To and The relevant features represent learnable parameters, participate in the calculation of attention weights, and work together with other parameters to determine the final attention weights; All of these are represented as learnable parameters.

[0028] The beneficial effects of adopting the above-mentioned further scheme are that the Re-WKV attention mechanism integrates a bidirectional attention mechanism, allowing the receptive field to cover the entire world, replacing the quadratic complexity of traditional self-attention with linear computational complexity, and also introducing a cyclic attention mechanism to apply bidirectional attention cyclically in different directions such as horizontal and vertical, enhancing global token interaction and more effectively capturing 2D image dependencies.

[0029] Based on the above technical solution, the present invention can be further improved as follows.

[0030] As a further improvement to this technical solution, the Re-WKV attention mechanism also applies bidirectional attention cyclically along different scanning directions (horizontal and vertical scans) through a cyclic attention mechanism, as shown in the following formula: ;

[0031] in Indicates the first the attention result obtained in the m-th loop; the first bidirectional WKV attention calculation operation; the m-th bidirectional WKV attention calculation operation; the direction change operation, the key-value pair the result after the direction change operation; the first bidirectional WKV attention calculation operation; the attention result obtained in the m-th loop the direction change operation;

[0032] the initial value the output is ; after m loops, the final Re-WKV attention result is obtained;

[0033] and the wkv is combined with the result after the Sigmoid processing through element-level multiplication to obtain the output ; specifically, the Sigmoid maps to 0-1 to adjust the weight of the attention result; then, through element-level multiplication, the attention result is combined with the adjusted to obtain the final output ; and is combined with the input feature through residual connection, and a scaling factor is introduced;

[0034] The Channel Mix part also first passes through the Layer Norm and the omnidirectional token shift layer to generate and two branches, and then is activated through the Squared ReLU to obtain the output , and then is combined with the result after the Sigmoid processing through element-level multiplication to obtain the output ; a residual block is introduced to combine with the input feature through residual connection, and a scaling factor is introduced.

[0035] The beneficial effect of adopting the above further scheme is that the application realizes efficient fusion of key features of low-light RAW images by virtue of the unique CRF module design. Local features, i.e., the close connection and characteristics of pixels in a small region of an image, such as the gray level change of the edge pixels of a small object, the texture direction, and other details, the convolution branch of the application adopts a 3x3 convolution layer to capture short-range pixel dependency with a limited receptive field, accurately extract local features and suppress noise; global features are related to the overall structure and layout of an image, covering long-range dependencies in different regions, such as the spatial position distribution of objects, the overall light and dark contrast trend, etc. The RWKV branch incorporates innovative omnidirectional token shift layers and Re-WKV modules to achieve linear complexity of calculation while maintaining the ability to model long-range dependencies with a multi-branch structure and bidirectional attention mechanism, efficiently extracting global features. Through the CRF module, the two are successfully combined, solving the fusion deficiency problem of the prior art, significantly improving the low-light RAW image enhancement effect, and making the enhanced image clearly present local details while maintaining the overall structure intact and reasonable.

[0036] On the basis of the above technical solution, the application can also be improved as follows.

[0037] As a further improvement of the technical solution, in the multi-level encoder-decoder in S4, each encoder CRF is cascaded by a down-sampling layer and an up-sampling layer for the decoder's CRF;

[0038] The down-sampling layer in the encoder is composed of a 3x3 convolution layer and a shuffle operation; the up-sampling layer in the decoder is a transpose convolution; the encoder and the decoder are spliced at the same layer feature and the decoder up-sampling feature, and input to the CTF block for feature fusion to restore the size of the feature map.

[0039] As a further improvement of the technical solution, S5 maps the feature map at the end of the decoder to after a 3x3 convolution layer , using a LeakyReLU activation function after the convolution layer: , then shuffling to render the enhanced RGB image .

[0040] The beneficial effect of adopting the above further scheme is that the application directly enhances the RAW format image, avoiding the loss of detailed information in the data conversion process of traditional methods based on RGB images, and can better exploit and utilize the rich information in RAW data, thereby realizing higher quality image enhancement effect.

[0041] In addition to the purposes, features and advantages described above, the application has other purposes, features and advantages. The application will be further described below with reference to the drawings. BRIEF DESCRIPTION OF DRAWINGS

[0042] Figure 1 The overall working procedure flowchart of the present application;

[0043] Figure 2 The working principle flowchart of S3 of the present application;

[0044] Figure 3 The working principle flowchart of the omni-directional token displacement layer in S3 of the present application;

[0045] Figure 4 The working principle flowchart of S3 of the present application. DETAILED DESCRIPTION

[0046] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the protection scope of the present application.

[0047] A low-light RAW image enhancement method based on RawRWKV in an embodiment of the present application includes the following steps:

[0048] S1: perceiving the input low-light RAW image in the camera RGGB sensor, i.e., the single-channel RGGB Bayer array (the single-channel RGB Bayer array is specifically in the original data state of the low-light RAW), using the MCR and SID data sets to train the model, in order to further increase the diversity of the training samples, using the data enhancement technology to randomly crop, randomly horizontally flip and vertically flip the channel image, and adjusting the image size to 512x512 as the input;

[0049] When testing the model, a low-light RAW image is used for sampling, the size of the input image is HxWx1, and the low-light RAW image is converted into a multi-channel input, and the corresponding expression is as follows:

[0050] perceiving the low-light RAW image , the RAW image corresponding size ; set as the output channel image, the size of the output channel image is , and the calculation formula is: ; wherein is the output image coordinate, is the input channel index (i corresponding to ); indicates the RAW image In the middle, the coordinates are , the pixel value of the channel index 0.

[0051] S2: Extract the features of each channel image, and use the activation function to increase the nonlinearity of the model. The specific operation formula is: , wherein is the channel image, is the convolution layer for extracting local features, is the activation function, is the channel image feature;

[0052] The convolution embedding layer slides the convolution kernel over the channel image , covering the local area of the channel image , then multiplies each element in the convolution kernel with the pixel value of the corresponding position of the image, and adds the product to obtain the data. Then repeat the above operation at each position of the channel image to obtain the local feature, and the corresponding expression is as follows:

[0053] ;

[0054] , wherein is the pixel value of the input image at coordinates and channel , is the weight value of the convolution at the corresponding position, is the bias value of the channel ;

[0055] , wherein the convolution kernel weight is a learnable parameter that is constantly adjusted during the training process to extract information that best represents the features of the image. For example, the convolution kernel weight captures local features such as edges and textures in the image. For multi-channel image input, the convolution kernel performs convolution operation on each channel image and combines the results.

[0056] If there is only a convolution layer, the model is linear and can only learn the linear relationship between the input and the output (such as multiplying all pixel brightness by a coefficient). It cannot solve complex problems in low-light images, such as the need for stronger brightness enhancement for dark features and the need to maintain details (non-linear enhancement) for bright parts; the distribution difference between noise and signal (more noise in the dark part, which needs to be suppressed and amplified; strong signal in the bright part, which needs to be preserved) and other problems. Therefore, nonlinearity is introduced into the model through , and the corresponding expression is: ;

[0057] , wherein is the set function; specifically, , , the function is linearly activated with normal slope, retaining strong signal features; When, , is the slope and is positive, the slope is activated, and the dark features are amplified with a slope .

[0058] S3, build a CRF module, please refer to the structure shown in Figure 2 : the CRF module includes two parallel branches, convolution branch and RWKV branch; integrate the results of the two branches, output the feature map through two convolution layers and an activation layer;

[0059] Further, in the S3 step, the convolution branch is used to obtain the local features of the image, and the corresponding expression is ; the RWKV branch structure includes Spatial Mix and Channel Mix two parts;

[0060] The Spatial Mix part adopts Layer Norm to normalize the local features, specifically normalizing the mean of the local features to 0 and the variance to 1, and the corresponding expression is: the perception input feature matrix is , where is the number of samples, is the feature dimension; the normalization calculation is performed through the formula: , where the mean of the local features in the dimension is ; the variance of the local features in the dimension is ; is a constant; and are learnable parameters, respectively used for scaling and shifting the normalized features to adapt to the learning needs of the model; is the element in the row and the column of the feature matrix .

[0061] In order to solve the deficiency of the existing omnidirectional token shift layer in capturing the local dependence of 2D images, the existing omnidirectional token shift layer can only shift tokens from limited directions, and cannot fully utilize the correlation of adjacent tokens in each direction in 2D images. Specifically, Spatial Mix uses an omnidirectional token shift layer (structure please refer to Figure 3 ) to adjust the spatial position information, and the omnidirectional token shift layer combines different size convolution operations (such as ) and the input itself to weight and combine the features, and the corresponding expression is:

[0062] ;

[0063] wherein is the output of Layer Norm, i.e., the normalized local features; is a weight coefficient, used to control the contribution proportion of different scale convolution operations and the input itself in the omnidirectional token shift layer; respectively represent the convolution operations of , which shift and fuse tokens from various directions; different branches are responsible for token shifting in different context ranges, and this multi-branch design can more accurately capture local context information, thereby enhancing the model's perception of local features in low-light RAW images;

[0064] In the test phase, in order to reduce the computational cost and improve efficiency, the omnidirectional token shift layer adopts a structure reparameterization strategy, which utilizes the feature that small convolution kernel weights can be integrated into larger convolution kernel weights through zero padding. The multi-branch structure during training is merged into a single-branch structure with a convolution kernel size of 5. This ensures the accuracy of the model's capture of local dependencies during testing, while also improving the running efficiency. As a result, the low-light RAW image enhancement method based on RawRWKV can better restore local details and enhance the clarity and quality of images when processing low-light RAW images, effectively improving the visual effect of low-light images.

[0065] The omnidirectional token shift layer generates three branches to extract and adjust spatial position information from different scales, enhancing the model's understanding of image spatial structure, wherein ; wherein are three linear projection layers, which linearly transform to obtain three branches, respectively;

[0066] After the omnidirectional token shift layer operation, and are input to Re-WKV, and output is obtained through the Re-WKV attention mechanism, and the corresponding expression is: ; ;

[0067] wherein is the data that integrates all local image feature unit information, represents the attention weight value corresponding to the th local image feature unit in ; represents the total number of local image feature units; is used to traverse 1 to Index variables for local image feature unit information; for The Middle The key value corresponding to each local image feature unit. To and The relevant features represent learnable parameters, participate in the calculation of attention weights, and work together with other parameters to determine the final attention weights; All are represented as learnable parameters, used to adjust the method and intensity of attention calculation, so that the attention result for each token is... The receptive field is determined by all other tokens, ensuring it covers the entire spectrum from the first token to the last. This avoids the quadratic computational complexity caused by query-key matrix multiplication in standard self-attention mechanisms, resulting in a computational complexity of O(n). ,in It's the number of tokens. It refers to the number of channels, which allows the model to acquire global information when processing 2D images.

[0068] Simultaneously, the Re-WKV attention mechanism also applies bidirectional attention cyclically along different scanning directions (H-Scan and V-Scan) through a cyclic attention mechanism, as detailed in the following formula: ;

[0069] in Indicates the first The attention result obtained from the next iteration; For the first This involves a bidirectional WKV attention computation operation; To change direction, For key values The result after the direction change operation; To make the first Attention results obtained in the second iteration Perform a direction change operation; and the initial value The output is After M iterations, the final Re-WKV attention result is obtained. This approach allows the model to capture information from images from different directions, enhancing the interaction between global tokens and more effectively modeling dependencies in 2D images;

[0070] Furthermore, in practical applications, due to the number of loops... Much smaller than the total number of tokens The Re-WKV attention mechanism enhances global token interaction capabilities while maintaining consistency with... The same linear computational complexity improves model performance without increasing computational burden, enabling the RawRWKV-based low-light RAW image enhancement method to more efficiently restore image details and improve image quality when processing low-light RAW images.

[0071] Further, wkv is combined with The result after Sigmoid processing is obtained by element-level multiplication to get the output ; Specifically, Sigmoid maps between 0 and 1 to adjust the weight of the attention result, so that the model can pay more attention to important features; then through element-level multiplication, the attention result is combined with the adjusted to get the final output , the corresponding expression is: ;

[0072] where represents Sigmoid function processing on , mapping its value to between 0 and 1; represents element-level multiplication, i.e., element-level multiplication of the Sigmoid result and ; represents a linear projection layer, which is used to further linearly transform the multiplied result to get the final output ;

[0073] In deep neural networks, as the number of network layers increases, the gradient may gradually become smaller when passing to the previous layers during backpropagation, and even tend to zero, which is the gradient vanishing problem. When the gradient vanishes, the parameter update of the previous layers becomes very slow, or even almost impossible to update, making it difficult for the model to train and learn effective feature representations. Therefore, to alleviate the problem of gradient vanishing, the is combined with the input features through residual connection; residual connection adds a shortcut in the network, so that the input directly skips some layers and is added to the output of the later layers, so that during backpropagation, the gradient is directly passed to the previous layers through the shortcut without going through layer-by-layer multiplication operations, making the parameter update of the previous layers. The corresponding expression is: ;

[0074] To scale attention, a scaling factor is introduced, which can adjust the attention result so that the model can more flexibly focus on different features according to the actual situation;

[0075] In the low-light image enhancement model, by scaling attention, the model can more accurately capture important features in the image. For example, in a low-light environment, the edges and textures of objects in the image may be blurred. By adjusting the scaling factor, the model can pay more attention to these detailed features, thereby better enhancing the image and improving its quality and clarity.

[0076] The Channel Mix part in the RWKV branch structure is also processed by Layer Norm and the omnidirectional token shift layer to generate and two branches, and is activated by Squared ReLU to obtain the output , and then the result after Sigmoid processing of is multiplied element-wise to obtain the output .

[0077] S4, a multi-level encoder-decoder based on CRF is constructed to generate a corresponding feature map Each encoder CRF module is cascaded by a down-sampling layer and an up-sampling layer for the decoder CRF module.

[0078] Further, in the S4 step, a 4-level encoder-decoder architecture is adopted to effectively enhance the low-light RAW image by gradually extracting and fusing image features. The encoder is responsible for down-sampling and feature extraction of the input image, gradually reducing the size of the feature map and increasing the number of channels to capture high-level semantic information of the image. The decoder, on the other hand, restores the size of the image through up-sampling and feature fusion operations and generates the final enhanced image.

[0079] The down-sampling layer in the encoder consists of a 3x3 convolution layer and a shuffle operation. The 3x3 convolution layer is used to reduce the number of channels by transforming the input features through convolution operations to extract more representative features while reducing the dimensionality of the data and reducing the amount of computation. The shuffle operation is used to reduce the size of the features by rearranging the positions of the pixels without losing information, and the corresponding expression is: the perception output feature map is , and the output feature map after pixel shuffling is , where is the shuffle factor.

[0080] The up-sampling layer in the decoder is a transposed convolution that expands the size of the feature map to restore the original size of the image. By performing interpolation and convolution operations on the feature map, the resolution of the feature map is increased, and the corresponding expression is: the perception input feature map is , and the transposed convolution kernel is , is the kernel size; bias is , then the output feature map is ;

[0081] The encoder same-level features are spliced with the decoder up-sampling features, input to the CTF block for feature fusion, to restore the size of the feature map, to obtain the output .

[0082] S5, construct an output module, apply a convolutional layer and a shuffle operation, output an enhanced RGB image.

[0083] Further, in the S5 step, the decoder end feature map is mapped to by a 3x3 convolutional layer , and the 3x3 convolutional layer further transforms the features, adjusting the number of channels of the feature map to 12, preparing for the subsequent shuffle operation and generating an RGB image. Through convolutional operation, higher-level features can be extracted, enabling the model to better understand the content of the image and generate more accurate enhanced images;

[0084] A LeakyReLU activation function is used after the convolutional layer, i.e. , then a shuffle is performed to rearrange the pixels in the feature map, converting the previously processed feature map into the format of an RGB image, rendering an enhanced RGB image .

[0085] Further, in the above steps, the low-light RAW image enhancement model uses python language to write code and pytorch framework, using short-exposure images in the MCR and SID datasets as input, and their corresponding long-exposure images as target output. The dataset contains a large number of images under low-light conditions and corresponding high-quality images. By training on these datasets, the model can learn the mapping relationship between low-light images and high-quality images, thereby achieving enhancement of low-light images. The model uses L1 loss to measure the difference between the predicted output and the target output, and the calculation formula of L1 loss is: ; where is the number of samples, is the predicted output of the model, is the target output;

[0086] The optimizer uses Adam, which is an adaptive learning rate optimization algorithm that automatically adjusts the learning rate based on the gradient of the parameters, improving the training efficiency and stability of the model; the learning rate adjustment uses a cosine annealing strategy (initial learning rate ), which is a method of dynamically adjusting the learning rate according to the number of training rounds, gradually reducing the learning rate in the form of a cosine function

[0087] Further, for low-light image processing, the RAW image that needs to be processed for image enhancement is input into the RawRWKV image enhancement model, and the enhanced image is output. The results show that the model can well recover the texture and details from the low-light image, achieving the best visual effect, and even the naked eye is difficult to distinguish the difference between the predicted output and the true result.

[0088] RawRWKV is compared with the current most advanced CNN-based models (SID, SGN, LDC and ABSID) and Transformer-based models (Uformer and Restormer). Since the models Uformer and Restormer are not based on RAW data processing, the above models are re-implemented, and the original 3-channel RGB input layer is changed so that these models can accept 4-channel RAW RGGB data without shuffling. These models are trained using the same strategy as RawRWKV; the results of the CNN-based models come from their corresponding papers. The evaluation indicators of the model results are PSNR (peak signal-to-noise ratio) and SSIM (structural similarity index).

[0089] The above shows and describes the basic principles, main features and advantages of the present application. Those skilled in the art should understand that the present application is not limited to the above embodiments, and the above embodiments and descriptions in the specification are only preferred examples of the present application and are not intended to limit the present application. Without departing from the spirit and scope of the present application, various changes and improvements can be made to the present application, and these changes and improvements all fall within the scope of the claimed present application. The scope of protection of the present application is defined by the appended claims and their equivalents.

Claims

1. A low-light RAW image enhancement method based on RWKV, characterized in that, The method comprises the following steps: S1, adjusting the size and state of the RAW image by using data enhancement technology, and expanding the training data set; Using the original data as the original data input for the test set; sampling the input image in the test set and converting it into a multi-channel input; S2, extracting local features of the multi-channel input through a convolution embedding layer, and using an activation function to increase the nonlinearity of the model; S3, constructing a CRF module, including a parallel convolution branch and a RWKV branch, integrating the results of the two branches, and outputting the feature map obtained through two convolution layers and an activation layer; The convolution branch in S3 is used to obtain local features of the image; The RWKV branch structure includes Spatial Mix and Channel Mix two parts; The Spatial Mix part uses Layer Norm to normalize the local features; The Spatial Mix uses an omnidirectional token shift layer to adjust the spatial position information, and the omnidirectional token shift layer shifts and fuses tokens from all directions; different branches are responsible for token shifting in different context ranges; S4, constructing a multi-level encoder-decoder based on CRF to generate a corresponding feature map; S5, constructing an output module, applying a convolution layer and a shuffle operation, and outputting an enhanced RGB image.

2. The RWKV-based low-light RAW image enhancement method according to claim 1, characterized in that: The expression of sampling in S1 is as follows: perceptual low-light RAW image , RAW image Corresponding size , and are the height and width of the RAW image respectively; set as the output channel image, the size of the output channel image is , then the calculation formula is: ; wherein is the output image coordinate, is the input channel index; represents the pixel value of the RAW image , the coordinate is , and the channel index is 0.

3. The RWKV-based low-light RAW image enhancement method according to claim 1, characterized in that: The convolutional embedding layer in S2 is through convolution kernel Slide across the channel image to cover the channel image. The local region is then processed, and each element in the convolution kernel is multiplied by the corresponding image pixel value. The products are then summed to obtain the data. This process is repeated at each position in the channel image to obtain local features.

4. The RWKV-based low-light RAW image enhancement method according to claim 3, characterized in that: The expression of the increase in model nonlinearity in S2 is: ; wherein is a set function; specifically, , , the function is linearly activated with normal slope, preserving strong signal features; , , is a slope and is positive, the activation is amplified with slope for dark features with slope .

5. The RWKV-based low-light RAW image enhancement method according to claim 1, characterized in that: The convolution branch in the S3 is used to obtain the local features of the image, and the corresponding expression is: The RWKV branch structure includes two parts of Spatial Mix and Channel Mix. The Spatial Mix part adopts Layer Norm, normalizes the local features, and specifically normalizes the mean of the local features to 0 and the variance to 1. The corresponding expression is: the perceptual input feature matrix is wherein is the number of samples, is the feature dimension; The normalization calculation is performed by the formula: where the mean of the local feature in dimension is and the variance of the local feature in dimension is ; is a constant; and are learnable parameters; is the element in the th row and th column of the feature matrix .

6. The RWKV-based low-light RAW image enhancement method according to claim 5, characterized in that: The omnidirectional token shift layer combines convolution operations of different sizes and the input itself to weight and combine features, and its corresponding expression is: ; wherein is the output of Layer Norm, i.e., the normalized local features; is a weight coefficient, used to control the contribution proportion of different scale convolution operations and the input itself in the omnidirectional token shift layer; respectively represent the convolution operations of .

7. The RWKV-based low-light RAW image enhancement method according to claim 6, characterized in that: The omni-directional token shift layer generates three branches, extract and adjust spatial position information from different scales, wherein ; wherein is a three linear projection layer, by linear projection layer on linear transformation, respectively, get three branches; After the operation of the omnidirectional token shift layer, the output is input to the Re-WKV, and the Re-WKV attention mechanism is used to obtain and the output, and the corresponding expression is: ; ; ; in This is data that integrates information from all local image feature units. Represented as The Middle Attention weight values ​​corresponding to each local image feature unit This represents the total number of feature units in a local image. For traversing 1 to Index variables for local image feature unit information; for The Middle The key value corresponding to each local image feature unit. To and The relevant features represent learnable parameters, participate in the calculation of attention weights, and work together with other parameters to determine the final attention weights; All of these are represented as learnable parameters.

8. The RWKV-based low-light RAW image enhancement method according to claim 7, characterized in that: The Re-WKV attention mechanism also applies bidirectional attention along horizontal scanning and vertical scanning in different scanning directions through a cyclic attention mechanism, and the specific formula is: ; wherein denotes the attention result of the th loop computation; is the th bidirectional WKV attention computation operation; is the direction change operation, is the key-value result after the direction change operation; is the attention result of the th loop computation is subjected to the direction change operation; Initial value The input is ; after M cycles, the final Re-WKV attention result ; and wkv The result after Sigmoid processing is multiplied by element level to get the output ; The specific Sigmoid maps the attention result between 0-1 to adjust the weight; and then through element-level multiplication, the attention result is combined with the adjusted to get the final output ; and the input feature is connected in residual connection, and a scaling factor is introduced; The Channel Mix part is also processed by Layer Norm and omni-directional token shift layer first to generate and two branches, and then is activated by Squared ReLU to obtain the output , and then the result after Sigmoid processing of is multiplied element by element to obtain the output ; the residual block is introduced to connect the residual of and the input feature, and a scaling factor is introduced.

9. The RWKV-based low-light RAW image enhancement method according to claim 2, characterized in that: In the multi-level encoder-decoder in S4, each encoder CRF is cascaded by a down-sampling layer and an up-sampling layer for the decoder CRF; The down-sampling layer in the encoder is composed of a 3x3 convolution layer and a shuffle operation; The up-sampling layer in the decoder is a transpose convolution; the same layer features of the encoder and the up-sampling features of the decoder are spliced and input into the CTF block for feature fusion to restore the size of the feature map.

10. The RWKV-based low-light RAW image enhancement method according to claim 9, characterized in that: The S5 maps the decoder-end feature map to a 3x3 convolutional layer , followed by a LeakyReLU activation function , then shuffled, rendering the enhanced RGB image .

Citation Information

Patent Citations

  • Meteorological forecasting method of MIM-rwkv improved by SADBO based on big data framework

    CN117388953A

  • Global and local multi-scale fused infrared guide low-light image enhancement method

    CN119399045A