Low-quality image correction method and device, storage medium and electronic equipment
By identifying and processing low-light and rainy images, the problem of poor image quality in different driving environments is solved, high-quality image correction in low-light and rainy conditions is achieved, and the safety of autonomous driving technology is ensured.
Patent Information
- Application Number
- CN202510519728.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-24
AI Technical Summary
In different driving environments, how to ensure the quality of images collected by image sensors, especially in low light and rainy conditions.
By identifying the type of the target image, light-enhancing processing is performed for low-light images, and rain-removing processing is performed for raining images, thereby obtaining high-quality corrected images.
In different driving environments, by separately processing low-light and rain images, the clarity and resolution of the image are significantly improved, ensuring the safety of autonomous driving technology.
Smart Images

Figure CN120047363A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and more particularly, to a method, apparatus, storage medium, and electronic device for correcting low-quality images. Background Art
[0002] In recent years, with the continuous development of technologies in the field of artificial intelligence, autonomous driving technology has become an important driving force for the development of the automotive industry. Currently, it is in the transition from assisted driving to conditional autonomous driving and also in the stage of starting highly autonomous driving, with broad development prospects.
[0003] Some current autonomous driving technologies rely on images collected by image sensors around the vehicle body for autonomous driving control. The clarity and distinguishability of the images directly affect the safety of autonomous driving. Therefore, how to ensure the quality of the collected images in different driving environments has become a difficult problem that concerns those skilled in the art. Summary of the Invention
[0004] The purpose of the present invention is to provide a method, apparatus, storage medium, and electronic device for correcting low-quality images to improve the above problems.
[0005] To achieve the above purpose, the technical solutions adopted in the embodiments of the present invention are as follows: In a first aspect, an embodiment of the present invention provides a method for correcting low-quality images, the method comprising: Identifying the type of the target image; When the target image is a low-light image, performing light enhancement processing on the low-light image to obtain a low-light corrected image; When the target image is a rainy image, performing rain removal processing on the rainy image to obtain a rainy corrected image; wherein the low-light image is an image with a brightness lower than a brightness threshold and does not include rainy features, and the rainy image is an image including rainy features.
[0006] In a second aspect, an embodiment of the present invention provides a device for correcting low-quality images, the device comprising: A first processing unit for identifying the type of the target image; A second processing unit for, when the target image is a low-light image, performing light enhancement processing on the low-light image to obtain a low-light corrected image; The second processing unit is further configured to, when the target image is a rainy image, perform rain removal processing on the rainy image to obtain a rainy corrected image; wherein the low-light image is an image with a brightness lower than a brightness threshold and does not include rainy features, and the rainy image is an image including rainy features.
[0007] In a third aspect, an embodiment of the present invention provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above method is implemented.
[0008] In a fourth aspect, an embodiment of the present invention provides an electronic device, the electronic device includes: a processor and a memory, the memory is used to store one or more programs; when the one or more programs are executed by the processor, the above method is implemented.
[0009] Compared with the prior art, a low-quality image correction method, device, storage medium and electronic device provided by an embodiment of the present invention include: identifying the type of a target image; when the target image is a low-light image, performing light enhancement processing on the low-light image to obtain a low-light corrected image; when the target image is a rainy image, performing rain removal processing on the rainy image to obtain a rainy corrected image; wherein, the low-light image is an image with a brightness lower than a brightness threshold and does not include rainy features, and the rainy image is an image including rainy features. By identifying the type of the target image and separately correcting the rainy image and the low-light image, the quality of images collected in different driving environments is ensured.
[0010] In order to make the above objects, features and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, the detailed description is as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0012] Figure 1 It is a schematic structural diagram of the electronic device provided by an embodiment of the present invention.
[0013] Figure 2 It is a schematic flowchart of the low-quality image correction method provided by an embodiment of the present invention.
[0014] Figure 3 It is a schematic diagram of the units of the low-quality image correction device provided by an embodiment of the present invention.
[0015] In the figure: 10 - processor; 11 - memory; 12 - bus; 13 - communication interface; 401 - first processing unit; 402 - second processing unit. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. The components of the embodiments of the present invention generally described and illustrated in the figures herein can be arranged and designed in a variety of different configurations.
[0017] Therefore, the detailed description of the embodiments of the present invention provided in the drawings below is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0018] It should be noted that like reference numerals and letters denote like items in the following figures, and thus, once an item is defined in one figure, it need not be further defined and explained in subsequent figures. At the same time, in the description of the present invention, the terms "first", "second", etc. are only used for descriptive distinction and cannot be construed as indicating or implying relative importance.
[0019] It should be noted that, in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variation thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0020] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "upper", "lower", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings or the orientation or positional relationship in which the inventive product is customarily placed during use. It is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.
[0021] In the description of the present invention, it should also be noted that, unless otherwise clearly specified and defined, the terms "arrangement" and "connection" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected, or indirectly connected through an intermediate medium, and it may be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0022] The following will describe in detail some embodiments of the present invention with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0023] Embodiments of the present invention provide an electronic device, which may be a vehicle computer, a server device, a mobile phone device, etc. Please refer to Figure 1 , the structural schematic diagram of the electronic device. The electronic device includes a processor 10, a memory 11, and a bus 12. The processor 10 and the memory 11 are connected through the bus 12, and the processor 10 is used to execute an executable module stored in the memory 11, such as a computer program.
[0024] The processor 10 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the low-quality image correction method may be completed by the integrated logic circuit in the hardware of the processor 10 or an instruction in software form. The above-mentioned processor 10 may be a general-purpose processor, including a central processing unit (Central Processing Unit, abbreviated as CPU), a network processor (Network Processor, abbreviated as NP), etc.; it may also be a digital signal processor (Digital Signal Processor, abbreviated as DSP), an application specific integrated circuit (Application Specific Integrated Circuit, abbreviated as ASIC), a field programmable gate array (Field-Programmable Gate Array, abbreviated as FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0025] The memory 11 may include a high-speed random access memory (RAM: Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory.
[0026] The bus 12 may be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. Figure 1 Although only one bidirectional arrow is used in the figure, it does not mean that there is only one bus 12 or only one type of bus 12 .
[0027] The memory 11 is used to store programs, such as programs corresponding to the low-quality image correction device. The low-quality image correction device includes at least one software function module that can be stored in the memory 11 in the form of software or firmware or fixed in the operating system (OS) of the electronic device. After receiving the execution instruction, the processor 10 executes the program to implement the low-quality image correction method.
[0028] Possibly, the electronic device provided by the embodiment of the present invention further includes a communication interface 13. The communication interface 13 is connected to the processor 10 via a bus.
[0029] It should be understood that Figure 1 The structure shown is only a schematic diagram of a portion of the electronic device. The electronic device may also include Figure 1 More or fewer components as shown, or with Figure 1 Different configurations shown. Figure 1 Each component shown in the figure can be implemented by hardware, software or a combination thereof.
[0030] A low-quality image correction method provided by an embodiment of the present invention can be applied to, but not limited to, Figure 1 For detailed procedures, please refer to the electronic equipment shown in Figure 2 , including: S10, S20 and S30, which are described in detail as follows.
[0031] S10, identifying the type of the target image. When the target image is a low-light image, executing S20; when the target image is a rainy image, executing S30.
[0032] The low-light image is an image with a brightness lower than a brightness threshold and does not include rain features, and the rain image is an image including rain features. The target image is an image around the car. In an optional implementation, cameras are arranged in front, rear, left front, left rear, right front and right rear of the car to collect images around the car in real time.
[0033] Optionally, the target image is input into a classifier for classification. The classifier consists of a convolutional neural network. After multiple layers of convolution and pooling, the features of the image will be gradually abstracted. Finally, through the fully connected layer, these abstracted features are mapped to three category labels: "normal image", "low-light image", and "rainy image" for subsequent targeted operations.
[0034] S20, perform light enhancement on the low-light image to obtain a low-light corrected image.
[0035] In an alternative embodiment, the low-light image collected in a low-light environment is sent to a light enhancement module based on natural language supervision for light enhancement processing to obtain a low-light corrected image.
[0036] S30, perform rain removal on the rainy image to obtain a rain-corrected image.
[0037] Optionally, the rainy image collected in a rainy environment is sent to a light enhancement module based on the Fourier learning strategy to obtain a rain-corrected image.
[0038] In the low-quality image correction method provided by the embodiments of the present invention, by identifying the type of the target image, the rainy image and the low-light image are corrected separately to ensure the quality of the images collected in different driving environments.
[0039] In Figure 2 On this basis, regarding the content in S20, the embodiments of the present invention also provide an alternative embodiment. Please refer to the following. S20, perform light enhancement on the low-light image to obtain a low-light corrected image, including: S21, S22, S23, and S24, which are specifically described as follows.
[0040] S21, input the low-light image into a 3×3 convolutional layer to extract shallow features .
[0041] Among them, , , H represents the length of the image, W represents the width of the image, and 3 and C are the number of channels corresponding to the image or features.
[0042] S22, use the feature enhancement unit to process the shallow features .
[0043] Among them, the feature enhancement unit includes multiple feature enhancement modules. The input feature of the first feature enhancement module is the shallow feature , and the input feature of the i-th feature enhancement module is the output feature of the (i - 1)-th feature enhancement module, where 2 ≤ i ≤ the total number of feature enhancement modules in the feature enhancement unit.
[0044] In an alternative embodiment, the feature enhancement unit includes 3 feature enhancement modules, and the output of the first feature enhancement module is , which will be input into the second feature enhancement module to obtain , which will be input into the third feature enhancement module to obtain .
[0045] Optionally, the feature enhancement module includes an information fusion attention part and a natural language supervision part. The information fusion attention part includes a channel attention block, a pixel attention block, and a cross-layer attention fusion block. On this basis, the embodiments of the present invention also provide an alternative embodiment on how the feature enhancement module works. Please refer to the following. S22, using the feature enhancement unit to process the shallow features , including: S221, S222, S223, S224, S225, and S226, which are specifically described as follows.
[0046] S221, processing the input feature of the s-th feature enhancement module through the channel attention block to obtain the output of the channel attention block of the s-th feature enhancement module.
[0047] Optionally, S221, processing the input feature of the s-th feature enhancement module through the channel attention block to obtain the output of the channel attention block of the s-th feature enhancement module, including: S221-1, S221-2, and S221-3, which are specifically as follows.
[0048] S221-1, performing global average pooling on the input feature of the s-th feature enhancement module to obtain the first global average pooling result .
[0049]
[0050] Among them, represents the first global average pooling result (including the global features of each channel), H represents the length of the input feature (also the length of the image), W represents the width of the input feature (also the width of the image), represents the pixel value of the pixel point with coordinates (m, n) in the input feature .
[0051] S221-2, according to the first global average pooling result , determine the feature channel weights corresponding to the s-th feature enhancement module .
[0052] Optionally, in order to determine the weights of each channel, the first global average pooling result will pass through two convolutional layers, and then be activated using the Sigmoid and ReLU functions to obtain the feature channel weights .
[0053]
[0054] Among them, represents the feature channel weights corresponding to the s-th feature enhancement module.
[0055] S221-3, according to the input features and the feature channel weights corresponding to the s-th feature enhancement module , determine the output of the channel attention block of the s-th feature enhancement module .
[0056] Optionally, the output of the channel attention block of the s-th feature enhancement module , and then obtain the output after enhancing important channel features and suppressing unimportant channel features.
[0057] S222, process the output of the channel attention block of the s-th feature enhancement module through the pixel attention block to obtain the output of the pixel attention block of the s-th feature enhancement module .
[0058] Optionally, S222, process the output of the channel attention block of the s-th feature enhancement module through the pixel attention block to obtain the output of the pixel attention block of the s-th feature enhancement module , including S222-1, S222-2, and S222-3, which are specifically described as follows.
[0059] S222-1, according to the output of the channel attention block , obtain the feature pixel weights corresponding to the s-th feature enhancement module .
[0060] Optionally, the pixel attention layer first inputs the output of the channel attention block into two convolutional layers, and then uses the Sigmoid and ReLU functions to activate, obtaining the feature pixel weights corresponding to the s-th feature enhancement module .
[0061]
[0062] Among them, represents the feature pixel weight corresponding to the s-th feature enhancement module.
[0063] S222-2. According to the feature pixel weight corresponding to the s-th feature enhancement module and the output of the channel attention block , determine the output of the pixel attention block of the s-th feature enhancement module .
[0064] The output after enhancing the feature information of the low-light area and the high-frequency area can be obtained. , .
[0065] S223. Use the cross-layer attention fusion block to process the input feature , the output of the channel attention block and the residual connection result of the input feature , as well as the output of the pixel attention block and the residual connection result of the input feature to generate the output of the cross-layer attention fusion block of the s-th feature enhancement module .
[0066] Optionally, S223. Use the cross-layer attention fusion block to process the input feature , the output of the channel attention block and the residual connection result of the input feature , as well as the output of the pixel attention block and the residual connection result of the input feature to generate the output of the cross-layer attention fusion block of the s-th feature enhancement module , including: S223-1, S223-2, S223-3, S223-4, and S223-5, which are specifically described as follows.
[0067] S223-1. Stack the input feature , the output of the channel attention block and the residual connection result of the input feature , as well as the output of the pixel attention block and the residual connection result of the input feature before and after to obtain the reshaping result .
[0068] Among them, .
[0069] S223-2. Use a 1×1 convolutional layer to integrate the reshaping result to obtain the integrated information.
[0070] S223-3, based on the integrated information, generate the first query matrix using a 3×3 depth convolution , the first key matrix and the first value matrix .
[0071] S223-4, according to the first query matrix , the first key matrix and the first value matrix , obtain the cross-layer attention weighted feature.
[0072]
[0073]
[0074] Among them, represents the cross-layer attention weighted feature, is a 3×3 layer-related attention matrix, is the scale factor, and perform two-dimensional reshaping on the first query matrix , the first key matrix and the first value matrix respectively to obtain two-dimensional , and , two-dimensional , , .
[0075] S223-5, according to the cross-layer attention weighted feature and the reshaping result , determine the output of the cross-layer attention fusion block of the s-th feature enhancement module.
[0076]
[0077] Among them, represents the output of the cross-layer attention fusion block of the s-th feature enhancement module.
[0078] S224, divide the output of the cross-layer attention fusion block into n levels to obtain the image feature .
[0079] Among them, represents the i-th level feature, which can be used as the image feature input for the natural language supervision part.
[0080] S225, input the low-light image into the text feature extraction model to obtain the low-light image Text features 。
[0081] Among them, represents the i-th text feature, which is used to describe the text of the image. The text feature extraction model can be, but is not limited to, a pre-trained CLIP model.
[0082] S226, using the cross-attention mechanism as the fusion layer, for the text features and image features are fused to obtain the output features of the feature enhancement module.
[0083] Optionally, S226, using the cross-attention mechanism as the fusion layer, for the text features and image features are fused to obtain the output features of the feature enhancement module, including: S226-1, S226-2, and S226-3, which are specifically described as follows.
[0084] S226-1, converting the text features through the encoder into the vector form of 。
[0085] It should be understood that the vector form of is more convenient to process.
[0086] S226-2, based on the vector form of and image features generate the second query matrix , the second key matrix and the second value matrix 。
[0087] Among them, , , , among which, represents the projection matrix corresponding to the query matrix, represents the projection matrix corresponding to the key matrix, represents the projection matrix corresponding to the value matrix, represents the intermediate representation form of the image features 。
[0088] S226-3, obtaining the attention-weighted feature according to the second query matrix , the second key matrix and the second value matrix 。
[0089]
[0090] Among them, B is a pre-configured relative position encoding matrix, is the scaling factor.
[0091] It should be noted that the attention-weighted feature is the output of the natural language supervision part and also the final output of the feature enhancement module.
[0092] S23, fuse the output features of each feature enhancement module to obtain the first fused feature .
[0093] Taking the number of feature enhancement modules as 3 as an example, finally, the outputs , , from the three feature enhancement modules are input into the feature fusion module. Arrange them front and back to form .
[0094] S24, process the first fused feature using two layers of 3×3 convolutional layers, and input the processing result into a 1×1 convolutional layer to output the low-light corrected image.
[0095] The final output number of channels is 3, and the corrected low-light corrected image .
[0096] Based on Figure 2 , regarding the content in S30, the embodiment of the present invention also provides an optional implementation manner. Please refer to the following. S30, perform rain removal processing on the rainy image to obtain the rain-corrected image, including: S31, S32, and S33, which are specifically described as follows.
[0097] S31, input the rainy image into a 3×3 convolutional layer to extract shallow features .
[0098] Among them, .
[0099] S32, input the shallow features into the multi-scale U-Net architecture to obtain deeper features.
[0100] Among them, the multi-scale U-Net architecture includes multiple groups of Fourier residual state space blocks. The input feature of the first group of Fourier residual state space blocks is the shallow feature , and the input feature of the i-th group of Fourier residual state space blocks is the output feature of the (i - 1)-th group of Fourier residual state space blocks, where 2 ≤ i ≤ the total number of Fourier residual state space blocks.
[0101] In an alternative embodiment, each group of Fourier residual state space blocks includes a Fourier space interaction state space model and a Fourier channel evolution state space model. On this basis, with regard to the content in S32, the embodiments of the present invention also provide an alternative embodiment. Please refer to the following text. S32, input the shallow features into the multi-scale U-Net architecture to obtain deeper features, including: S32A~S32K, which are specifically described as follows.
[0102] S32A, in the Fourier space interaction state space model, use the normalization technique (LayerNorm) to normalize the input features to generate normalized features .
[0103] By performing normalization, the problem of gradient vanishing or explosion can be reduced.
[0104] S32B, in the Fourier space interaction state space model, adopt the collaborative processing of the Fourier branch and the spatial branch to obtain the output of the Fourier branch and the output of the spatial branch ; Optionally, S32B, in the Fourier space interaction state space model, adopt the collaborative processing of the Fourier branch and the spatial branch to obtain the output of the Fourier branch and the output of the spatial branch , including: S32B1 - S32B5, which are specifically described as follows.
[0105] S32B1, in the Fourier branch, perform a fast Fourier transform on the normalized features to transform the normalized features from the current spatial domain to the Fourier space to obtain the amplitude spectrum and the phase spectrum P .
[0106] Among them, the amplitude spectrum and the phase spectrum P are the representations of the normalized features in the frequency domain. The amplitude spectrum represents the intensity information of the frequency, and the phase spectrum P represents the phase information of the frequency.
[0107] It should be noted that in order to promote the interaction between spatial information and frequency information, the collaborative processing of the Fourier branch and the spatial branch is adopted.
[0108] S32B2, use the zigzag encoding to scan and rearrange the amplitude spectrum and the phase spectrum P respectively to obtain a one-dimensional amplitude sequence and one-dimensional phase sequence .
[0109] In order to orderly associate the relationships between different frequencies, scanning and rearrangement are required. Among them, the one-dimensional amplitude sequence and one-dimensional phase sequence are arranged in ascending order of frequency.
[0110] S32B3 sends the one-dimensional amplitude sequence and one-dimensional phase sequence to the frequency Mamba block to obtain the output of the Fourier branch .
[0111] Optionally, S32B3 sends the one-dimensional amplitude sequence and one-dimensional phase sequence to the frequency Mamba block to obtain the output of the Fourier branch , including: S32B3-1, S32B3-2, S32B3-3, and S32B3-4, which are specifically described as follows.
[0112] S32B3-1 passes the one-dimensional amplitude sequence and one-dimensional phase sequence sequentially through a depthwise separable convolutional layer → SiLU activation function → SSM framework → LayerNorm normalization layer to obtain the first amplitude intermediate product and the first phase intermediate product .
[0113] Thus, the dynamic relationships between different frequencies are fully captured, with low-frequency components enhanced and high-frequency components suppressed.
[0114] S32B3-2 performs an inverse fast Fourier transform on the first amplitude intermediate product and the first phase intermediate product to obtain the second fused feature .
[0115] By performing an inverse fast Fourier transform, the feature is converted back to the current spatial domain.
[0116] S32B3-3 inputs the normalized feature into the activation function SiLU to obtain .
[0117] S32B3-4 performs an element-wise multiplication of the second fused feature with to obtain the output of the Fourier branch .
[0118] By multiplying the second fused feature with Perform element-wise multiplication to dynamically adjust the importance of frequency components, better remove raindrops and restore background details, and finally obtain the output of the Fourier branch .
[0119] S32B4. In the spatial branch, perform a 1×1 convolution operation on the normalized features to extract features between channels.
[0120] S32B5. Feed the features between channels into the spatial Mamba block to obtain the output of the spatial branch .
[0121] ; where represents the output of the spatial branch, represents the output of the spatial Mamba block.
[0122] S32C. According to the output of the Fourier branch and the output of the spatial branch , obtain the output of the Fourier-spatial interaction state space model .
[0123] Optionally, stack the output of the Fourier branch and the output of the spatial branch front and back, and perform a 1×1 convolution operation on the stacked result. The number of output channels is , to obtain the output of the Fourier-spatial interaction state space model , which is also used as the input of the Fourier-channel evolution state space model .
[0124] S32D. Perform global average pooling on the input of the Fourier-channel evolution state space model to obtain the second global average pooling result .
[0125]
[0126] where represents the second global average pooling result (which effectively encapsulates the global information of the features), represents the length of represents the width of represents the pixel value of the pixel point with coordinates (m, n) in the feature .
[0127] S32E. On the second global average pooling result Perform a channel Fourier transform to obtain the channel Fourier transform result 。
[0128]
[0129] Among them, is the number of channels, z is the index in the frequency domain, j is the imaginary unit, is the normalized value of the frequency, represents the value of the c-th channel at a certain spatial position.
[0130] S32F, according to the real and imaginary parts in the channel Fourier transform result obtain the amplitude component and the phase component.
[0131]
[0132]
[0133] Among them, represents the amplitude component, represents the phase component, represents the channel Fourier transform result in the real part, represents the channel Fourier transform result in the imaginary part.
[0134] S32G, pass the amplitude component and the phase component sequentially through the depthwise separable convolutional layer → SiLU activation function → SSM framework → LayerNorm normalization layer to obtain the second amplitude intermediate product and the second phase intermediate product 。
[0135] Enhance the low-frequency components and suppress the high-frequency components to obtain a feature sequence with stronger representation ability 、 。
[0136] S32H, use the inverse channel Fourier transform to convert the second amplitude intermediate product and the second phase intermediate product back to the current spatial domain to obtain the third fusion feature.
[0137] S32I, input the input of the Fourier channel evolution state space model into the activation function SiLU to obtain 。
[0138] S32J, combine the third fusion feature with Perform an element-wise multiplication to obtain the output of the Fourier channel evolution state space model .
[0139] By element-wise multiplying the third fused feature with perform an element-wise multiplication to dynamically adjust the importance of frequency components, better remove raindrops and restore background details, and obtain .
[0140] S32K, multiply the output of the Fourier space interaction state space model with the output of the Fourier channel evolution state space model to obtain the output of the Fourier residual state space block .
[0141]
[0142] wherein, represents the output of the Fourier residual state space block.
[0143] S33, input the output of the last Fourier residual state space block into a 1×1 convolutional layer to output a rain-corrected image with 3 output channels.
[0144] After several similar Fourier residual state space blocks and finally passing through a 1×1 convolutional layer with 3 output channels, the corrected image is finally obtained.
[0145] Please refer to Figure 3 , Figure 3 for a low-quality image correction device provided by an embodiment of the present invention. Optionally, the low-quality image correction device is applied to the electronic device described above.
[0146] The low-quality image correction device includes: a first processing unit 401 and a second processing unit 402.
[0147] The first processing unit 401 is used to identify the type of the target image; The second processing unit 402 is used to perform light enhancement on the low-light image to obtain a low-light corrected image when the target image is a low-light image; The second processing unit 402 is further used to perform rain removal on the rain image to obtain a rain-corrected image when the target image is a rain image; wherein, the low-light image is an image with a brightness lower than the brightness threshold and does not include rain features, and the rain image is an image including rain features.
[0148] It should be noted that the low-quality image correction device provided in this embodiment can execute the method flow shown in the above method flow embodiment to achieve the corresponding technical effects. For a brief description, for the parts not mentioned in this embodiment, reference may be made to the corresponding content in the above embodiments.
[0149] An embodiment of the present invention also provides a storage medium that stores computer instructions and programs. When the computer instructions and programs are read and run, they execute the low-quality image correction method of the above embodiment. The storage medium may include memory, flash memory, registers, or a combination thereof, etc.
[0150] The following provides an electronic device, which may be a vehicle computer, a server device, a mobile phone device, etc. As shown in Figure 1 it, the above-mentioned low-quality image correction method can be implemented; specifically, the electronic device includes: a processor 10, a memory 11, and a bus 12. The processor 10 may be a CPU. The memory 11 is used to store one or more programs. When the one or more programs are executed by the processor 10, the low-quality image correction method of the above embodiment is executed.
[0151] In summary, a low-quality image correction method, device, storage medium, and electronic device provided by an embodiment of the present invention include: identifying the type of the target image; when the target image is a low-light image, performing light enhancement processing on the low-light image to obtain a low-light corrected image; when the target image is a rainy image, performing rain removal processing on the rainy image to obtain a rainy corrected image; wherein, the low-light image is an image with a brightness lower than the brightness threshold and does not include rainy features, and the rainy image is an image including rainy features. By identifying the type of the target image and separately correcting the rainy image and the low-light image, the quality of the images collected in different driving environments is ensured.
[0152] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0153] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to include all changes falling within the meaning and scope of the equivalent elements of the claims in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.
Claims
1. A method for correcting low-quality images, characterized in that: The method comprises: Identify the type of target image; When the target image is a low-light image, performing light enhancement processing on the low-light image to obtain a low-light corrected image; When the target image is a rainy image, performing rain removal processing on the rainy image to obtain a rain-corrected image; The low-light image is an image whose brightness is lower than a brightness threshold and does not include rain features, and the rain image is an image that includes rain features.
2. The low-quality image correction method according to claim 1, characterized in that: The step of performing light enhancement processing on the low-light image to obtain a low-light corrected image includes: The low light image Input to the 3×3 convolution layer to extract shallow features ; Using feature enhancement units to enhance shallow features Processing; wherein the feature enhancement unit includes multiple feature enhancement modules, and the input feature of the first feature enhancement module is a shallow feature , the input feature of the i-th feature enhancement module is the output feature of the i-1-th feature enhancement module, 2≤i≤the total number of feature enhancement modules in the feature enhancement unit; The output features of each feature enhancement module are fused to obtain the first fused feature ; Use two 3×3 convolutional layers to fusion the first feature The processed image is fed into a 1×1 convolutional layer to output a low-light corrected image.
3. The low-quality image correction method according to claim 2, characterized in that: The feature enhancement unit is used to enhance the shallow features Processing includes: The input features of the sth feature enhancement module are processed through the channel attention block Process to get the output of the channel attention block of the sth feature enhancement module ; The output of the channel attention block of the sth feature enhancement module is processed by the pixel attention block Processed to obtain the output of the pixel attention block of the sth feature enhancement module ; Use cross-layer attention fusion blocks to fused input features , the output of the channel attention block With input features The residual connection result and the output of the pixel attention block With input features The residual connection result of is processed to generate the output of the cross-layer attention fusion block of the s-th feature enhancement module ; The output of the cross-layer attention fusion block Divided into n levels to obtain image features ; The low light image Input to text feature extraction model for low-light images Text features ; Using the cross attention mechanism as the fusion layer, the text features and image features The fusion is performed to obtain the output features of the feature enhancement module.
4. The low-quality image correction method according to claim 3, characterized in that: The input features of the sth feature enhancement module are processed by the channel attention block Process to get the output of the channel attention block of the sth feature enhancement module ,include: Input features of the sth feature enhancement module Perform global average pooling to obtain the first global average pooling result ; in, represents the first global average pooling result, and H represents the input feature The length of W represents the input feature The width, Represents input features The pixel value of the pixel point with coordinates (m, n); According to the first global average pooling result , determine the feature channel weight corresponding to the sth feature enhancement module ; in, Represents the feature channel weight corresponding to the s-th feature enhancement module; According to the input characteristics The feature channel weight corresponding to the s-th feature enhancement module , determine the output of the channel attention block of the sth feature enhancement module .
5. The low-quality image correction method according to claim 3, characterized in that: The cross attention mechanism is used as the fusion layer to analyze the text features. and image features Fusion is performed to obtain the output features of the feature enhancement module, including: Text features Through the encoder Convert to vector form ; In vector form and image features Based on this, the second query matrix is generated , the second key matrix And the second value matrix ; According to the second query matrix , the second key matrix And the second value matrix Get attention weighted features ; Among them, B is the pre-configured relative position encoding matrix, is the scaling factor, attention weighted feature As the output of the feature enhancement module.
6. The low-quality image correction method according to claim 1, characterized in that: The step of performing rain removal processing on the rainy image to obtain a rain-corrected image includes: It will rain images Input to the 3×3 convolution layer to extract shallow features ; The shallow features Input the multi-scale U-Net architecture to obtain deeper features; the multi-scale U-Net architecture includes multiple groups of Fourier residual state space blocks, and the input features of the first group of Fourier residual state space blocks are shallow features. , the input features of the i-th group of Fourier residual state space blocks are the output features of the i-1-th group of Fourier residual state space blocks, 2≤i≤the total number of Fourier residual state space blocks; The output of the last Fourier residual state-space block , input 1×1 convolution layer, and output rain-corrected image with 3 channels.
7. The low-quality image correction method according to claim 6, characterized in that: The shallow features Input the multi-scale U-Net architecture to obtain deeper features, including: In the Fourier space interactive state space model, the input features are normalized using normalization techniques to generate normalized features ; In the Fourier space interactive state space model, the Fourier branch and the spatial branch are used for collaborative processing to obtain the output of the Fourier branch and the output of the spatial branch ; According to the output of the Fourier branch and the output of the spatial branch , and the output of the Fourier space interactive state space model is obtained ; Input to the Fourier channel evolution state-space model Perform global average pooling to obtain the second global average pooling result ; According to the channel Fourier transform results The real and imaginary parts of the signal are used to obtain the amplitude and phase components; The amplitude component and phase component are respectively passed through the depth-separable convolution layer → SiLU activation function → SSM framework → LayerNorm normalization layer to obtain the second amplitude intermediate product and the second phase intermediate ; Using the inverse channel Fourier transform, the second amplitude intermediate product and the second phase intermediate Convert back to the current spatial domain to obtain the third fusion feature; The input of the Fourier channel evolution state space model Input the activation function SiLU to get ; The third fusion feature is combined with Perform element-wise multiplication to obtain the output of the Fourier channel evolution state space model ; The output of the Fourier space interactive state space model Output of the Fourier channel evolution state space model Multiply them together to get the output of the Fourier residual state space block .
8. A low-quality image correction device, characterized in that: The device comprises: A first processing unit, used for identifying the type of the target image; a second processing unit, configured to, when the target image is a low-light image, perform light enhancement processing on the low-light image to obtain a low-light corrected image; The second processing unit is further configured to perform rain removal processing on the rain image to obtain a rain-corrected image when the target image is a rain image; The low-light image is an image whose brightness is lower than a brightness threshold and does not include rain features, and the rain image is an image that includes rain features.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. An electronic device, characterized in that: include: A processor and a memory, the memory being used to store one or more programs; When the one or more programs are executed by the processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Highway visor inclination identification method based on deep semantic segmentation network and image correction
CN114155518A
Low-light image enhancement method suitable for haze climate
CN119130878A
Different-emissivity infrared imaging correction method and system based on image enhancement technology
CN119251070A
System and processor implemented method for improved image quality and generating an image of a target illuminated by quantum particles
US20140340570A1