A low-quality image correction method, device, storage medium and electronic device
By identifying and targeted correction of low-light and rainy images, the problem that image acquisition quality affects the safety of autonomous driving is solved, and image quality is improved to ensure the safety of autonomous driving.
Patent Information
- Application Number
- CN202510519728.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-24
AI Technical Summary
In different driving environments, the clarity and resolution of images affect the safety of autonomous driving, and the prior art is difficult to ensure the quality of image acquisition.
By identifying the type of the target image, light-enhancing processing is performed for low-light images, rain-removing processing is performed on rainy images, and correction is performed separately to ensure image quality.
In different driving environments, improve the quality of image acquisition and ensure the safety of autonomous driving.
Smart Images

Figure CN120047363B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular, to a method, apparatus, storage medium, and electronic device for correcting low-quality images. Background Art
[0002] In recent years, with the continuous development of technologies in the field of artificial intelligence, autonomous driving technology has become an important driving force for the development of the automotive industry. Currently, it is in the transition from assisted driving to conditional autonomous driving, and at the same time, it is in the stage of the start of highly autonomous driving, with broad development prospects.
[0003] Some current autonomous driving technologies rely on images collected by image sensors around the vehicle body for autonomous driving control. The clarity and resolvability of the images directly affect the safety of autonomous driving. Therefore, how to ensure the quality of the collected images in different driving environments has become a difficult problem that concerns those skilled in the art. Summary of the Invention
[0004] The purpose of the present invention is to provide a method, apparatus, storage medium, and electronic device for correcting low-quality images to improve the above problems.
[0005] To achieve the above purpose, the technical solutions adopted in the embodiments of the present invention are as follows:
[0006] In a first aspect, an embodiment of the present invention provides a method for correcting low-quality images, the method including:
[0007] Identifying the type of a target image;
[0008] When the target image is a low-light image, performing light enhancement processing on the low-light image to obtain a low-light corrected image;
[0009] When the target image is a rainy image, performing rain removal processing on the rainy image to obtain a rainy corrected image;
[0010] Wherein, the low-light image is an image with a brightness lower than a brightness threshold and does not include rainy features, and the rainy image is an image including rainy features.
[0011] In a second aspect, an embodiment of the present invention provides a device for correcting low-quality images, the device including:
[0012] A first processing unit for identifying the type of a target image;
[0013] A second processing unit for, when the target image is a low-light image, performing light enhancement processing on the low-light image to obtain a low-light corrected image;
[0014] The second processing unit is further configured to perform rain removal processing on the rain image when the target image is a rain image, so as to obtain a rain-corrected image;
[0015] Wherein, the low-light image is an image with a brightness lower than a brightness threshold and does not include rain features, and the rain image is an image including rain features.
[0016] In a third aspect, an embodiment of the present invention provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above method is implemented.
[0017] In a fourth aspect, an embodiment of the present invention provides an electronic device, the electronic device includes: a processor and a memory, the memory is used to store one or more programs; when the one or more programs are executed by the processor, the above method is implemented.
[0018] Compared with the prior art, a low-quality image correction method, device, storage medium and electronic device provided by an embodiment of the present invention include: identifying the type of the target image; when the target image is a low-light image, performing light enhancement processing on the low-light image to obtain a low-light corrected image; when the target image is a rain image, performing rain removal processing on the rain image to obtain a rain-corrected image; wherein, the low-light image is an image with a brightness lower than a brightness threshold and does not include rain features, and the rain image is an image including rain features. By identifying the type of the target image and separately correcting the rain image and the low-light image, the quality of the images collected in different driving environments is guaranteed.
[0019] In order to make the above objects, features and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, the detailed description is as follows. Description of the Drawings
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0021] Figure 1 It is a schematic structural diagram of the electronic device provided by an embodiment of the present invention.
[0022] Figure 2 It is a schematic flowchart of the low-quality image correction method provided by an embodiment of the present invention.
[0023] Figure 3 It is a schematic diagram of the units of the low-quality image correction device provided by an embodiment of the present invention.
[0024] In the figure: 10 - processor; 11 - memory; 12 - bus; 13 - communication interface; 401 - first processing unit; 402 - second processing unit. Detailed implementation manner
[0025] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. Generally, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0026] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0027] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0028] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variant thereof are intended to cover non - exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0029] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "upper", "lower", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of the present invention is usually placed during use. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.
[0030] In the description of the present invention, it should also be noted that unless otherwise clearly specified and defined, the terms "arrangement" and "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0031] The following will describe in detail some embodiments of the present invention with reference to the drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0032] The embodiments of the present invention provide an electronic device, which can be an on-board computer, a server device, a mobile phone device, etc. Please refer to Figure 1 , the structural schematic diagram of the electronic device. The electronic device includes a processor 10, a memory 11, and a bus 12. The processor 10 and the memory 11 are connected through the bus 12, and the processor 10 is used to execute the executable module stored in the memory 11, such as a computer program.
[0033] The processor 10 can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the low-quality image correction method can be completed by the integrated logic circuit in the hardware of the processor 10 or the instructions in software form. The above-mentioned processor 10 can be a general-purpose processor, including a central processing unit (Central Processing Unit, abbreviated as CPU), a network processor (Network Processor, abbreviated as NP), etc.; it can also be a digital signal processor (Digital Signal Processor, abbreviated as DSP), an application specific integrated circuit (Application Specific Integrated Circuit, abbreviated as ASIC), a field programmable gate array (Field-Programmable Gate Array, abbreviated as FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0034] The memory 11 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory.
[0035] The bus 12 may be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. Figure 1 Although only one bidirectional arrow is used in the figure, it does not mean that there is only one bus 12 or only one type of bus 12 .
[0036] The memory 11 is used to store programs, such as programs corresponding to the low-quality image correction device. The low-quality image correction device includes at least one software function module that can be stored in the memory 11 in the form of software or firmware or fixed in the operating system (OS) of the electronic device. After receiving the execution instruction, the processor 10 executes the program to implement the low-quality image correction method.
[0037] Possibly, the electronic device provided by the embodiment of the present invention further includes a communication interface 13. The communication interface 13 is connected to the processor 10 via a bus.
[0038] It should be understood that Figure 1 The structure shown is only a schematic diagram of a portion of the electronic device. The electronic device may also include Figure 1 More or fewer components as shown, or with Figure 1 Different configurations shown. Figure 1 Each component shown in the figure can be implemented by hardware, software or a combination thereof.
[0039] A low-quality image correction method provided by an embodiment of the present invention can be applied to, but not limited to, Figure 1 For detailed procedures, please refer to the electronic equipment shown in Figure 2 , including: S10, S20 and S30, which are described in detail as follows.
[0040] S10, identifying the type of the target image. When the target image is a low-light image, executing S20; when the target image is a rainy image, executing S30.
[0041] Among them, the low-light image is an image with a brightness lower than the brightness threshold and does not include rain features, and the rain image is an image including rain features. The target image is an image around the vehicle. In an alternative embodiment, cameras are arranged in front of, behind, left front, left rear, right front, and right rear of the vehicle to collect images around the vehicle in real time.
[0042] Optionally, the target image is input into a classifier for classification. The classifier consists of a convolutional neural network. After multiple layers of convolution and pooling, the features of the image will be gradually abstracted, and finally, through the fully connected layer, these abstracted features are mapped to three category labels: "normal image", "low-light image", and "rain image" for subsequent targeted operations.
[0043] S20, perform light enhancement on the low-light image to obtain a low-light corrected image.
[0044] In an alternative embodiment, the low-light image collected in a low-light environment is sent to a light enhancement module based on natural language supervision for light enhancement processing to obtain a low-light corrected image.
[0045] S30, perform rain removal on the rain image to obtain a rain-corrected image.
[0046] Optionally, the rain image collected in a rainy environment is sent to a light enhancement module based on Fourier learning strategy to obtain a rain-corrected image.
[0047] In the low-quality image correction method provided by the embodiments of the present invention, by identifying the type of the target image, the rain image and the low-light image are corrected separately to ensure the quality of the images collected in different driving environments.
[0048] On the basis of Figure 2 Regarding the content in S20, the embodiments of the present invention also provide an alternative embodiment. Please refer to the following. S20, perform light enhancement on the low-light image to obtain a low-light corrected image, including: S21, S22, S23, and S24, which are specifically described as follows.
[0049] S21, input the low-light image into a 3×3 convolutional layer to extract shallow features .
[0050] Among them, , , H represents the length of the image, W represents the width of the image, and 3 and C are the number of channels corresponding to the image or features.
[0051] S22, use the feature enhancement unit to process the shallow features for processing.
[0052] Among them, the feature enhancement unit includes a plurality of feature enhancement modules, and the input feature of the first feature enhancement module is the shallow feature , and the input feature of the i-th feature enhancement module is the output feature of the (i - 1)-th feature enhancement module, where 2 ≤ i ≤ the total number of feature enhancement modules in the feature enhancement unit.
[0053] In an optional implementation manner, the feature enhancement unit includes 3 feature enhancement modules, and the output of the first feature enhancement module is , which will be input into the second feature enhancement module to obtain , which will be input into the third feature enhancement module to obtain .
[0054] Optionally, the feature enhancement module includes an information fusion attention part and a natural language supervision part. The information fusion attention part includes a channel attention block, a pixel attention block, and a cross-layer attention fusion block. On this basis, the embodiments of the present invention also provide an optional implementation manner on how the feature enhancement module works. Please refer to the following. S22, using the feature enhancement unit to process the shallow feature , including: S221, S222, S223, S224, S225, and S226, which are specifically described as follows.
[0055] S221, processing the input feature of the s-th feature enhancement module through the channel attention block to obtain the output of the channel attention block of the s-th feature enhancement module.
[0056] Optionally, S221, processing the input feature of the s-th feature enhancement module through the channel attention block to obtain the output of the channel attention block of the s-th feature enhancement module, including: S221-1, S221-2, and S221-3, which are specifically as follows.
[0057] S221-1, performing global average pooling processing on the input feature of the s-th feature enhancement module to obtain the first global average pooling result .
[0058]
[0059] Among them, represents the first global average pooling result (including the global features of each channel), H represents the length of the input feature (which is also the length of the image), and W represents the input feature The width (which is also the width of the image), represents the input feature of the pixel at coordinates (m, n) in
[0060] S221-2. According to the first global average pooling result , determine the feature channel weight corresponding to the s-th feature enhancement module .
[0061] Optionally, in order to determine the weights of each channel, the first global average pooling result will pass through two convolutional layers, and then be activated using the Sigmoid and ReLU functions to obtain the feature channel weight .
[0062]
[0063] Among them, represents the feature channel weight corresponding to the s-th feature enhancement module.
[0064] S221-3. According to the input feature and the feature channel weight corresponding to the s-th feature enhancement module , determine the output of the channel attention block of the s-th feature enhancement module .
[0065] Optionally, the output of the channel attention block of the s-th feature enhancement module , and then obtain the output after enhancing important channel features and suppressing unimportant channel features.
[0066] S222. Process the output of the channel attention block of the s-th feature enhancement module through the pixel attention block to obtain the output of the pixel attention block of the s-th feature enhancement module .
[0067] Optionally, S222. Process the output of the channel attention block of the s-th feature enhancement module through the pixel attention block to obtain the output of the pixel attention block of the s-th feature enhancement module , including S222-1, S222-2, and S222-3, which are specifically described as follows.
[0068] S222-1. According to the output of the channel attention block , obtain the feature pixel weight corresponding to the s-th feature enhancement module .
[0069] Optionally, the pixel attention layer first outputs the output of the channel attention block Input into two convolutional layers, then activated using the Sigmoid and ReLU functions to obtain the feature pixel weights corresponding to the s-th feature enhancement module .
[0070]
[0071] Among them, represents the feature pixel weights corresponding to the s-th feature enhancement module.
[0072] S222-2, according to the feature pixel weights corresponding to the s-th feature enhancement module and the output of the channel attention block , determine the output of the pixel attention block of the s-th feature enhancement module .
[0073] The output after enhancing the feature information of the low-light area and the high-frequency area can be obtained, , .
[0074] S223, use the cross-layer attention fusion block to process the input feature , the output of the channel attention block and the residual connection result of the input feature , as well as the residual connection result of the output of the pixel attention block and the input feature to generate the output of the cross-layer attention fusion block of the s-th feature enhancement module .
[0075] Optionally, S223, use the cross-layer attention fusion block to process the input feature , the output of the channel attention block and the residual connection result of the input feature , as well as the residual connection result of the output of the pixel attention block and the input feature to generate the output of the cross-layer attention fusion block of the s-th feature enhancement module , including: S223-1, S223-2, S223-3, S223-4, and S223-5, which are specifically described as follows.
[0076] S223-1, stack the input feature , the output of the channel attention block and the residual connection result of the input feature , as well as the residual connection result of the output of the pixel attention block and the input feature front and back to obtain the reshaped result 。
[0077] Among them, 。
[0078] S223-2, use a 1×1 convolutional layer to integrate the reshaped result to obtain integrated information.
[0079] S223-3, based on the integrated information, use a 3×3 depth convolution to generate the first query matrix , the first key matrix and the first value matrix 。
[0080] S223-4, according to the first query matrix , the first key matrix and the first value matrix , obtain cross-layer attention weighted features.
[0081]
[0082]
[0083] Among them, represents the cross-layer attention weighted features, is a 3×3 layer-related attention matrix, is the scaling factor, and the first query matrix , the first key matrix and the first value matrix are respectively two-dimensionally reshaped to obtain two-dimensional , and , two-dimensional , , 。
[0084] S223-5, according to the cross-layer attention weighted features and the reshaped result , determine the output of the cross-layer attention fusion block of the s-th feature enhancement module.
[0085]
[0086] Among them, represents the output of the cross-layer attention fusion block of the s-th feature enhancement module.
[0087] S224, divide the output of the cross-layer attention fusion block into n levels to obtain the image features 。
[0088] Among them, represents the i-th level feature and can be used as the image feature input for the natural language supervision part.
[0089] S225, input the low-light image into the text feature extraction model to obtain the text features of the low-light image .
[0090] Among them, represents the i-th text feature, which is used to describe the text of the image. The text feature extraction model can be, but is not limited to, a pre-trained CLIP model.
[0091] S226, use the cross-attention mechanism as the fusion layer to fuse the text features and the image features to obtain the output features of the feature enhancement module.
[0092] Optionally, S226, use the cross-attention mechanism as the fusion layer to fuse the text features and the image features to obtain the output features of the feature enhancement module, including: S226-1, S226-2, and S226-3, which are specifically described as follows.
[0093] S226-1, convert the text features into the vector form through the encoder .
[0094] It should be understood that the vector form is more convenient to process.
[0095] S226-2, based on the vector form and the image features , generate the second query matrix , the second key matrix , and the second value matrix .
[0096] Among them, , , , among which, represents the projection matrix corresponding to the query matrix, represents the projection matrix corresponding to the key matrix, represents the projection matrix corresponding to the value matrix, represents the intermediate representation form of the image features .
[0097] S226-3, according to the second query matrix , the second key matrix and the second value matrix Obtain the attention-weighted feature .
[0098]
[0099] Among them, B is a pre-configured relative position encoding matrix, is the scaling factor.
[0100] It should be noted that the attention-weighted feature is the output of the natural language supervision part and also the final output of the feature enhancement module.
[0101] S23, fuse the output features of each feature enhancement module to obtain the first fused feature .
[0102] Taking the number of feature enhancement modules as 3 as an example, finally, the outputs from the three feature enhancement modules , , are input into the feature fusion module. Arrange them in sequence to form .
[0103] S24, process the first fused feature using two layers of 3×3 convolutional layers, and input the processing result into a 1×1 convolutional layer to output the low-light corrected image.
[0104] The final output channel number is 3, and the corrected low-light corrected image .
[0105] Based on Figure 2 , regarding the content in S30, the embodiment of the present invention also provides an optional implementation manner. Please refer to the following. S30, perform rain removal processing on the rain image to obtain a rain-corrected image, including: S31, S32, and S33, which are specifically described as follows.
[0106] S31, input the rain image into a 3×3 convolutional layer to extract shallow features .
[0107] Among them, .
[0108] S32, input the shallow features into the multi-scale U-Net architecture to obtain deeper features.
[0109] Among them, the multi-scale U-Net architecture includes multiple groups of Fourier residual state space blocks. The input feature of the first group of Fourier residual state space blocks is the shallow feature , and the input feature of the i-th group of Fourier residual state space blocks is the output feature of the (i - 1)-th group of Fourier residual state space blocks, where 2 ≤ i ≤ the total number of Fourier residual state space blocks.
[0110] In an alternative embodiment, each group of Fourier residual state space blocks includes a Fourier space interaction state space model and a Fourier channel evolution state space model. On this basis, with regard to the content in S32, the embodiments of the present invention further provide an alternative embodiment. Please refer to the following text. S32, input the shallow feature into the multi-scale U-Net architecture to obtain deeper features, including: S32A~S32K, which are specifically described as follows.
[0111] S32A, in the Fourier space interaction state space model, use the normalization technique (LayerNorm) to normalize the input feature to generate a normalized feature .
[0112] By performing the normalization process, the problem of gradient vanishing or explosion can be reduced.
[0113] S32B, in the Fourier space interaction state space model, adopt the collaborative processing of the Fourier branch and the spatial branch to obtain the output of the Fourier branch and the output of the spatial branch ;
[0114] Optionally, S32B, in the Fourier space interaction state space model, adopt the collaborative processing of the Fourier branch and the spatial branch to obtain the output of the Fourier branch and the output of the spatial branch , including: S32B1 - S32B5, which are specifically described as follows.
[0115] S32B1, in the Fourier branch, perform a fast Fourier transform on the normalized feature to convert the normalized feature from the current spatial domain to the Fourier space to obtain the amplitude spectrum and the phase spectrum P .
[0116] Among them, the amplitude spectrum and the phase spectrum P are the representations of the normalized feature in the frequency domain. The amplitude spectrum represents the intensity information of the frequency, and the phase spectrum P represents the phase information of the frequency.
[0117] It should be noted that, in order to promote the interaction between spatial information and frequency information, Fourier branches and spatial branches are used for collaborative processing.
[0118] S32B2, use zigzag encoding to scan and rearrange the amplitude spectrum and the phase spectrum P respectively, to obtain a one-dimensional amplitude sequence and a one-dimensional phase sequence .
[0119] In order to orderly associate the relationships between different frequencies, scanning and rearrangement are required. Among them, the frequencies in the one-dimensional amplitude sequence and the one-dimensional phase sequence are arranged in ascending order.
[0120] S32B3, send the one-dimensional amplitude sequence and the one-dimensional phase sequence into the frequency Mamba block to obtain the output of the Fourier branch .
[0121] Optionally, S32B3, send the one-dimensional amplitude sequence and the one-dimensional phase sequence into the frequency Mamba block to obtain the output of the Fourier branch , including: S32B3-1, S32B3-2, S32B3-3 and S32B3-4, which are specifically described as follows.
[0122] S32B3-1, send the one-dimensional amplitude sequence and the one-dimensional phase sequence sequentially through a depthwise separable convolution layer → SiLU activation function → SSM framework → LayerNorm normalization layer to obtain a first amplitude intermediate product and a first phase intermediate product .
[0123] Thus, the dynamic relationships between different frequencies can be fully captured, with low-frequency components enhanced and high-frequency components suppressed.
[0124] S32B3-2, perform inverse fast Fourier transform on the first amplitude intermediate product and the first phase intermediate product to obtain a second fused feature .
[0125] By performing inverse fast Fourier transform, the features are converted back to the current spatial domain.
[0126] S32B3-3, the normalized features Input the SiLU activation function to obtain .
[0127] S32B3-4, perform an element-wise product of the second fused feature with to obtain the output of the Fourier branch .
[0128] By performing an element-wise product of the second fused feature with , dynamically adjust the importance of frequency components, better remove raindrops and restore background details, and finally obtain the output of the Fourier branch .
[0129] S32B4, in the spatial branch, perform a 1×1 convolution operation on the normalized feature to extract features between channels.
[0130] S32B5, send the features between channels into the spatial Mamba block to obtain the output of the spatial branch .
[0131] ;
[0132] wherein represents the output of the spatial branch, represents the output of the spatial Mamba block.
[0133] S32C, based on the output of the Fourier branch and the output of the spatial branch, obtain the output of the Fourier spatial interaction state space model.
[0134] Optionally, stack the output of the Fourier branch and the output of the spatial branch front to back, perform a 1×1 convolution operation on the result of the front-to-back stacking, with the number of output channels being , obtain the output of the Fourier spatial interaction state space model, which is also used as the input of the Fourier channel evolution state space model.
[0135] S32D, perform global average pooling on the input of the Fourier channel evolution state space model to obtain the second global average pooling result .
[0136]
[0137] wherein Represents the second global average pooling result (which effectively encapsulates the global information of the features), Represents the length of Represents the width of Represents the pixel value of the pixel point with coordinates (m, n) in the feature
[0138] S32E, performs channel Fourier transform on the second global average pooling result to obtain the channel Fourier transform result .
[0139]
[0140] Among them, is the number of channels, z is the index in the frequency domain, j is the imaginary unit, is the normalized value of the frequency, Represents the value of the c-th channel at a certain spatial position.
[0141] S32F, obtains the amplitude component and the phase component according to the real part and the imaginary part in the channel Fourier transform result .
[0142]
[0143]
[0144] Among them, represents the amplitude component, represents the phase component, represents the real part in the channel Fourier transform result , represents the imaginary part in the channel Fourier transform result .
[0145] S32G, sequentially passes the amplitude component and the phase component through the depthwise separable convolutional layer → SiLU activation function → SSM framework → LayerNorm normalization layer to obtain the second amplitude intermediate product and the second phase intermediate product .
[0146] Enhances the low-frequency components and suppresses the high-frequency components to obtain a feature sequence 、 with stronger representation ability.
[0147] S32H, uses the inverse channel Fourier transform to transform the second amplitude intermediate product and the second phase intermediate product Convert it back to the current spatial domain to obtain the third fusion feature.
[0148] S32I, the input of the Fourier channel evolution state space model Input the activation function SiLU to obtain .
[0149] S32J, multiply the third fusion feature element-wise with to obtain the output of the Fourier channel evolution state space model .
[0150] By multiplying the third fusion feature element-wise with , the importance of frequency components is dynamically adjusted, raindrops are better removed, and background details are restored to obtain .
[0151] S32K, the output of the Fourier space interaction state space model is multiplied by the output of the Fourier channel evolution state space model to obtain the output of the Fourier residual state space block .
[0152]
[0153] wherein, represents the output of the Fourier residual state space block.
[0154] S33, the output of the last Fourier residual state space block , is input into a 1×1 convolutional layer to output a rain-corrected image with 3 output channels.
[0155] After several similar Fourier residual state space blocks, and finally passing through a 1×1 convolutional layer with 3 output channels, the finally corrected image is obtained.
[0156] Please refer to Figure 3 , Figure 3 which is a low-quality image correction device provided by an embodiment of the present invention. Optionally, the low-quality image correction device is applied to the electronic device described above.
[0157] The low-quality image correction device includes: a first processing unit 401 and a second processing unit 402.
[0158] The first processing unit 401 is used to identify the type of the target image;
[0159] A second processing unit 402, configured to perform light enhancement processing on the low-light image to obtain a low-light corrected image when the target image is a low-light image;
[0160] The second processing unit 402 is further configured to perform rain removal processing on the rainy image to obtain a rainy corrected image when the target image is a rainy image;
[0161] Wherein, the low-light image is an image with a brightness lower than the brightness threshold and does not include rain features, and the rainy image is an image including rain features.
[0162] It should be noted that the low-quality image correction device provided in this embodiment can execute the method flow shown in the above method flow embodiment to achieve the corresponding technical effects. For the sake of brief description, for the parts not mentioned in this embodiment, reference may be made to the corresponding content in the above embodiment.
[0163] An embodiment of the present invention further provides a storage medium, which stores computer instructions and programs. When the computer instructions and programs are read and run, they execute the low-quality image correction method of the above embodiment. The storage medium may include memory, flash memory, registers or a combination thereof, etc.
[0164] The following provides an electronic device, which may be a vehicle computer, a server device, a mobile phone device, etc. As shown in Figure 1 It can implement the above low-quality image correction method; specifically, the electronic device includes: a processor 10, a memory 11, and a bus 12. The processor 10 may be a CPU. The memory 11 is used to store one or more programs. When the one or more programs are executed by the processor 10, the low-quality image correction method of the above embodiment is executed.
[0165] In summary, a low-quality image correction method, device, storage medium and electronic device provided by an embodiment of the present invention include: identifying the type of the target image; when the target image is a low-light image, performing light enhancement processing on the low-light image to obtain a low-light corrected image; when the target image is a rainy image, performing rain removal processing on the rainy image to obtain a rainy corrected image; wherein, the low-light image is an image with a brightness lower than the brightness threshold and does not include rain features, and the rainy image is an image including rain features. By identifying the type of the target image and separately correcting the rainy image and the low-light image, the quality of the images collected in different driving environments is guaranteed.
[0166] The above are only preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
[0167] It is obvious to those skilled in the art that the present invention is not limited to the details of the above-mentioned exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or essential features of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
Claims
1. A low-quality image correction method, characterized in that, The method includes: Identifying the type of the target image; When the target image is a low-light image, performing light enhancement processing on the low-light image to obtain a low-light corrected image; When the target image is a rainy image, performing rain removal processing on the rainy image to obtain a rainy corrected image; Wherein, the low-light image is an image with a brightness lower than a brightness threshold and does not include rainy features, and the rainy image is an image including rainy features; The performing light enhancement processing on the low-light image to obtain a low-light corrected image includes: Input the low-light image into a 3×3 convolutional layer to extract shallow features ; Process the shallow features using a feature enhancement unit wherein the feature enhancement unit includes a plurality of feature enhancement modules, and the input feature of the first feature enhancement module is the shallow feature , and the input feature of the i-th feature enhancement module is the output feature of the (i - 1)-th feature enhancement module, where 2 ≤ i ≤ the total number of feature enhancement modules in the feature enhancement unit; Fuse the output features of each feature enhancement module to obtain the first fused feature ; Process the first fused feature using two 3×3 convolutional layers and input the processing result into a 1×1 convolutional layer to output a low-light corrected image; The use of the feature enhancement unit for the shallow features is processed as follows: Process the input features of the s-th feature enhancement module through the channel attention block to obtain the output of the channel attention block of the s-th feature enhancement module ; Process the output of the channel attention block of the $s$-th feature enhancement module through the pixel attention block to obtain the output of the pixel attention block of the $s$-th feature enhancement module ; Process the input features using a cross-layer attention fusion block and the output of the channel attention block along with the residual connection result of the input features and the output of the pixel attention block along with the residual connection result of the input features to generate the output of the cross-layer attention fusion block of the s-th feature enhancement module ; The output of the cross-layer attention fusion block is divided into n levels to obtain image features ; Input the low-light image into the text feature extraction model to obtain the text features of the low-light image ; Using a cross-attention mechanism as a fusion layer to fuse text features and image features to obtain the output features of the feature enhancement module; The input features of the s-th feature enhancement module are processed through the channel attention block to obtain the output of the channel attention block of the s-th feature enhancement module , including: The input features of the s-th feature enhancement module are subjected to global average pooling to obtain the first global average pooling result ; Among them, represents the first global average pooling result, H represents the input feature length, W represents the width of the input feature ; represents the input feature the pixel value of the pixel point with coordinates (m, n) in Based on the first global average pooling result , determine the feature channel weights corresponding to the s-th feature enhancement module ; Among them, represents the feature channel weight corresponding to the s-th feature enhancement module; According to the input features and the feature channel weights corresponding to the s-th feature enhancement module , determine the output of the channel attention block of the s-th feature enhancement module ; The performing rain removal processing on the rainy image to obtain a rainy corrected image includes: Rainy image Input it into a 3×3 convolutional layer to extract shallow features ; Input the shallow features into the multi-scale U-Net architecture to obtain deeper features; among them, the multi-scale U-Net architecture includes multiple groups of Fourier residual state space blocks, and the input feature of the first group of Fourier residual state space blocks is the shallow feature , and the input feature of the i-th group of Fourier residual state space blocks is the output feature of the (i-1)-th group of Fourier residual state space blocks, where 2 ≤ i ≤ the total number of Fourier residual state space blocks; Output of the last Fourier residual state space block , input a 1×1 convolutional layer, and output a rain-corrected image with 3 output channels; The shallow features are input into a multi-scale U-Net architecture to obtain deeper features, including: In the Fourier space interaction state space model, a normalization technique is used to normalize the input features to generate normalized features ; In the Fourier space interaction state space model, the Fourier branch and the spatial branch are processed collaboratively to obtain the output of the Fourier branch and the output of the spatial branch ; According to the output of the Fourier branch and the output of the spatial branch , the output of the Fourier spatial interaction state space model is obtained ; Input to the Fourier channel evolution state space model Perform global average pooling to obtain the second global average pooling result ; Based on the real and imaginary parts in the channel Fourier transform result obtain the amplitude component and the phase component; The amplitude component and the phase component are respectively passed through a depthwise separable convolutional layer → SiLU activation function → SSM framework → LayerNorm normalization layer in sequence to obtain a second amplitude intermediate product and a second phase intermediate product ; Using an inverse channel Fourier transform, the second amplitude intermediate product and the second phase intermediate product are transformed back to the current spatial domain to obtain a third fusion feature; Input the input of the Fourier channel evolution state space model into the activation function SiLU to obtain ; Multiply the third fusion feature with element-wise to obtain the output of the Fourier channel evolution state space model ; Multiply the output of the Fourier space interaction state space model with the output of the Fourier channel evolution state space model to obtain the output of the Fourier residual state space block .
2. The low-quality image correction method according to claim 1, wherein Using the cross-attention mechanism as a fusion layer for the text features and the image features to perform fusion to obtain the output features of the feature enhancement module, including: Convert text features Through an encoder Into vector form ; Based on the vector form of and the image features generate a second query matrix a second key matrix and a second value matrix ; According to the second query matrix , the second key matrix , and the second value matrix to obtain the attention weighted feature ; where B is a pre-configured relative position encoding matrix, is a scaling factor, and the attention-weighted feature serves as the output of the feature enhancement module.
3. A low-quality image correction device, characterized in that, The apparatus includes: A first processing unit for identifying the type of the target image; A second processing unit for, when the target image is a low-light image, performing light enhancement processing on the low-light image to obtain a low-light corrected image; The second processing unit is further configured to, when the target image is a rainy image, perform rain removal processing on the rainy image to obtain a rainy corrected image; Wherein, the low-light image is an image with a brightness lower than a brightness threshold and does not include rainy features, and the rainy image is an image including rainy features; The performing light enhancement processing on the low-light image to obtain a low-light corrected image includes: Input the low-light image into a 3×3 convolutional layer to extract shallow features ; Process the shallow features using a feature enhancement unit ; wherein, the feature enhancement unit includes a plurality of feature enhancement modules, the input feature of the first feature enhancement module is the shallow feature , and the input feature of the i-th feature enhancement module is the output feature of the (i - 1)-th feature enhancement module, where 2 ≤ i ≤ the total number of feature enhancement modules in the feature enhancement unit; Fuse the output features of each feature enhancement module to obtain the first fused feature ; Use two 3×3 convolutional layers to process the first fused feature and input the processing result into a 1×1 convolutional layer to output a low-light corrected image; The processing of the shallow features by using the feature enhancement unit includes: Process the input features of the s-th feature enhancement module through the channel attention block to obtain the output of the channel attention block of the s-th feature enhancement module ; Process the output of the channel attention block of the $s$-th feature enhancement module through the pixel attention block to obtain the output of the pixel attention block of the $s$-th feature enhancement module ; Using a cross-layer attention fusion block for the input features , the output of the channel attention block and the residual connection result of the input features as well as the output of the pixel attention block and the residual connection result of the input features are processed to generate the output of the cross-layer attention fusion block of the s-th feature enhancement module ; The output of the cross-layer attention fusion block is divided into n levels to obtain image features ; Input the low-light image into the text feature extraction model to obtain the text features of the low-light image ; Using the cross-attention mechanism as a fusion layer to fuse the text features and the image features to obtain the output features of the feature enhancement module; The input features of the s-th feature enhancement module are processed through the channel attention block to obtain the output of the channel attention block of the s-th feature enhancement module , including: The input features to the s-th feature enhancement module are subjected to global average pooling to obtain the first global average pooling result ; Among them, represents the first global average pooling result, and H represents the input feature length, and W represents the width of the input feature ; represents the input feature the pixel value of the pixel point with coordinates (m, n) in According to the first global average pooling result , determine the feature channel weights corresponding to the s-th feature enhancement module ; Among them, represents the feature channel weight corresponding to the s-th feature enhancement module; According to the input features and the feature channel weights corresponding to the s-th feature enhancement module , determine the output of the channel attention block of the s-th feature enhancement module ; The performing rain removal processing on the rainy image to obtain a rainy corrected image includes: Rainy image Input it into a 3×3 convolutional layer to extract shallow features ; Input the shallow features into a multi-scale U-Net architecture to obtain deeper features. Among them, the multi-scale U-Net architecture includes multiple groups of Fourier residual state space blocks. The input feature of the first group of Fourier residual state space blocks is the shallow feature , and the input feature of the i-th group of Fourier residual state space blocks is the output feature of the (i - 1)-th group of Fourier residual state space blocks, where 2 ≤ i ≤ the total number of Fourier residual state space blocks; Output of the last Fourier residual state space block , input a 1×1 convolutional layer, and output a rain-corrected image with 3 output channels; The shallow features are input into a multi-scale U-Net architecture to obtain deeper features, including: In the Fourier space interaction state space model, a normalization technique is used to normalize the input features to generate normalized features ; In the Fourier space interaction state space model, the Fourier branch and the spatial branch are processed collaboratively to obtain the output of the Fourier branch and the output of the spatial branch ; According to the output of the Fourier branch and the output of the spatial branch , the output of the Fourier spatial interaction state space model is obtained ; Input to the Fourier channel evolution state space model Perform global average pooling to obtain the second global average pooling result ; Based on the real and imaginary parts in the channel Fourier transform result obtain the amplitude component and the phase component; The amplitude component and the phase component are respectively passed through a depthwise separable convolutional layer → SiLU activation function → SSM framework → LayerNorm normalization layer in sequence to obtain a second amplitude intermediate product and a second phase intermediate product ; Using the inverse channel Fourier transform, the second amplitude intermediate product and the second phase intermediate product are transformed back to the current spatial domain to obtain the third fusion feature; Input the evolution state space model of the Fourier channel into the input activation function SiLU to obtain ; Multiply the third fusion feature with element-wise to obtain the output of the Fourier channel evolution state space model ; Multiply the output of the Fourier space interaction state space model with the output of the Fourier channel evolution state space model to obtain the output of the Fourier residual state space block .
4. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1-2.
5. An electronic device, characterized in that, It includes: A processor and a memory, where the memory is used to store one or more programs; When the one or more programs are executed by the processor, the method according to any one of claims 1-2 is implemented.
Citation Information
Patent Citations
Low-light image enhancement method suitable for haze climate
CN119130878A
Different-emissivity infrared imaging correction method and system based on image enhancement technology
CN119251070A