Image de-raining method, device, equipment and medium based on deformable multi-head attention

By introducing a deformation multi-head attention mechanism into the image rain removal method, combined with the architecture of encoder and decoder, the problem of insufficient spatial perception and long-range semantic capture capabilities in the existing technology is solved, and an efficient image rain removal effect is achieved, which is suitable for multi-scene and real-time applications.

CN117689580BActive Publication Date: 2025-05-30VOYAH AUTOMOBILE TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311708065.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-11
Publication Date
2025-05-30
Estimated Expiration
2043-12-11

AI Technical Summary

Technical Problem

The existing image enhancement methods for harsh weather environments have insufficient spatial perception and long-range semantic capture capabilities, making it difficult to remove noise such as rain lines, raindrops, rain fog and other noises at the same time, and the rain removal effect and computing power consumption are difficult to balance, which limits real-time processing applications.

Method used

The image rain removal method based on deformation multi-head attention is adopted. Through the combination of encoder and decoder, the deformation Transformer model and convolutional decoder are used to realize the deep semantic feature extraction and rain removal operation of the image.

Benefits of technology

It significantly improves the image enhancement effect, takes into account the extraction and processing of local and global semantic features of the image, and improves the model's spatial perception and long-range relationship perception capabilities of image features, and is suitable for real-time image rain removal in multiple scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117689580B_ABST
    Figure CN117689580B_ABST
Patent Text Reader

Abstract

The present invention discloses an image de-raining method, device, equipment and medium based on deformable multi-head attention, which relates to the field of image processing. The method includes performing continuous downsampling processing on an image to be processed based on a preset encoder to obtain deep semantic features of the image to be processed; performing upsampling processing on the deep semantic features based on a preset decoder to implement the de-raining operation on the image to be processed; wherein, the encoder is composed of multiple deformable Transformer models connected in series, the deformable Transformer model is constructed according to deformable multi-head self-attention and downsampling, and the decoder is created based on convolution and transposed convolution. This application can improve the image enhancement effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and in particular to an image de-raining method, device, equipment and medium based on deformable multi-head attention. Background Art

[0002] Computer vision perception, especially image processing technology for outdoor scenes, has become one of the most important components in the current research fields of intelligent transportation and intelligent driving. Image de-raining for actual road conditions has become one of the technologies with the most rapid development and the widest application in the field of artificial intelligence such as intelligent driving. This technology can be adapted to various important visual processing tasks including object detection, semantic segmentation, etc., and has important research and application values. This technology aims to process the images captured in rainy environments through algorithm models to remove raindrops, rain lines, rain fog, etc. from the images, so as to output high-quality images without noise.

[0003] However, the existing image enhancement methods for harsh weather environments have the following problems: 1. The spatial perception ability and long-range semantic capture ability of the encoders of existing end-to-end de-raining algorithms are limited, resulting in most image de-raining algorithms being difficult to remove raindrops, rain lines, rain fog and other noises simultaneously; 2. It is difficult to balance the de-raining effect and computing power consumption. Some cutting-edge technologies can achieve better de-raining effects, but they all have problems of excessive memory occupation or excessive computational complexity in the model inference stage, resulting in these technologies being difficult to achieve real-time processing, which greatly limits their applications in actual scenarios, such as scenarios for real-time perception in intelligent driving, real-time capture and processing of intelligent transportation images, etc.; 3. The details of the restored images are missing. In scenarios such as heavy rain or raindrops appearing on the lens, the raindrop noises interfere too much with the original image information, often resulting in the distortion of the restored content of the details of the rainy parts of the restored images. Summary of the Invention

[0004] The present application provides an image de-raining method, device, equipment and medium based on deformable multi-head attention, which can improve the image enhancement effect.

[0005] In a first aspect, an embodiment of the present application provides an image de-raining method based on deformable multi-head attention. The image de-raining method based on deformable multi-head attention includes:

[0006] Performing continuous downsampling processing on the image to be processed based on a preset encoder to obtain the deep semantic features of the image to be processed;

[0007] Performing upsampling processing on the deep semantic features based on a preset decoder to implement the de-raining operation on the image to be processed;

[0008] Among them, the encoder is composed of multiple deformable Transformer models connected in series. The deformable Transformer model is constructed based on deformable multi-head self-attention and downsampling, and the decoder is created based on convolution and transposed convolution.

[0009] Combined with the first aspect, in an implementation manner,

[0010] The deformable multi-head self-attention includes a global deformable multi-head self-attention mechanism and a local deformable multi-head self-attention mechanism;

[0011] The downsampling is implemented based on the PatchMering module in Swin Transformer.

[0012] Combined with the first aspect, in an implementation manner, for the global deformable multi-head self-attention mechanism, the specific implementation method is:

[0013] Process the input tensor through multiple connected feature map deformation transformation models, so that the elements in the input tensor are redistributed according to semantic information to obtain the processed input tensor;

[0014] Average the processed input tensor in the spatial dimension into multiple windows of a set size, and use the standard multi-head attention mechanism model to process the processed input tensor in each window to obtain the output of the global deformable multi-head self-attention mechanism;

[0015] Among them, the feature map deformation transformation model is constructed based on the deformation sampling mechanism in the deformable convolution algorithm.

[0016] Combined with the first aspect, in an implementation manner, for the local deformable multi-head self-attention mechanism, the specific implementation method is:

[0017] Divide the input tensor into multiple windows of a set size in the spatial dimension, and process each window with the feature map deformation transformation model to obtain multiple deformed window regions;

[0018] Execute the standard multi-head attention mechanism model in each deformed window region to obtain multiple tensors;

[0019] Concatenate the obtained multiple tensors in the spatial dimension to obtain an output tensor of a set dimension.

[0020] Combined with the first aspect, in an implementation manner,

[0021] The feature map deformation transformation model is used to convert any input tensor into an output tensor of the same dimension;

[0022] For converting to an output tensor of the same dimension, the specific steps include:

[0023] Filter the input tensor using a convolution with a set step size, convolution kernel, and output dimension, and calculate the bias tensor;

[0024] For any position in the input tensor, determine the corresponding position in the output tensor to obtain the output tensor.

[0025] Combined with the first aspect, in one implementation,

[0026] In the encoder, the output of each deformable Transformer model is used as the input of the next deformable Transformer model, and the output of the last deformable Transformer model is used as the output of the decoder;

[0027] The decoder includes a set number of basic decoders connected in series, and each basic decoder is composed of a convolution and a transposed convolution connected in series;

[0028] The input of the first basic decoder in the decoder is the output of the last deformable Transformer model in the encoder, and the inputs of the other basic decoders include a first part of the input and a second part of the input;

[0029] The first part of the input is the output of the previous basic decoder, and the second part of the input is the output of the corresponding deformable Transformer model in the encoder;

[0030] The output of the last basic decoder in the decoder is the de-rained image corresponding to the image to be processed.

[0031] In a second aspect, the present application provides an image de-raining method based on deformable multi-head attention. The image de-raining method based on deformable multi-head attention includes:

[0032] Construct a deformable Transformer model based on deformable multi-head self-attention and downsampling, connect multiple deformable Transformer models in series to obtain an encoder, and create a decoder based on convolution and transposed convolution

[0033] Perform continuous downsampling on the image to be processed based on the encoder to obtain the deep semantic features of the image to be processed;

[0034] Perform upsampling on the deep semantic features based on the decoder to implement the de-raining operation on the image to be processed.

[0035] In a third aspect, the present application provides an image de-raining device based on deformable multi-head attention. The image de-raining device based on deformable multi-head attention includes:

[0036] An acquisition module, which is used to perform continuous downsampling processing on the image to be processed based on a preset encoder to obtain the depth semantic features of the image to be processed;

[0037] An execution module, which is used to perform upsampling processing on the depth semantic features based on a preset decoder to realize the rain removal operation of the image to be processed;

[0038] Wherein, the encoder is composed of a plurality of deformable Transformer models connected in series, the deformable Transformer model is constructed according to deformable multi-head self-attention and downsampling, and the decoder is created based on convolution and transposed convolution.

[0039] In a fourth aspect, the present application provides an image rain removal device based on deformable multi-head attention. The image rain removal device based on deformable multi-head attention includes a processor, a memory, and an image rain removal program based on deformable multi-head attention stored on the memory and executable by the processor. When the image rain removal program based on deformable multi-head attention is executed by the processor, the steps of the above-mentioned image rain removal method based on deformable multi-head attention are realized.

[0040] In a fifth aspect, the present application provides a computer-readable storage medium, on which an image rain removal program based on deformable multi-head attention is stored. When the image rain removal program based on deformable multi-head attention is executed by a processor, the steps of the above-mentioned image rain removal method based on deformable multi-head attention are realized.

[0041] The beneficial effects brought by the technical solutions provided by the embodiments of the present application include:

[0042] (1) The encoding and decoding scheme proposed in the present application can take into account the extraction and processing of local and global semantic features of the image. In the form of combining deformable sampling and multi-scale multi-head attention mechanism, it can efficiently enhance the model's spatial perception and long-range relationship perception of image features, and significantly improve the image enhancement effect;

[0043] (2) The deformable multi-head self-attention mechanism proposed in the present application can be extended to the current mainstream feature learning models. At the cost of increasing the limited model complexity, it can significantly improve the model's feature learning ability, and is easily extended to deep models for downstream tasks such as image enhancement and image segmentation, and has good promotion and application value;

[0044] (3) The model architecture proposed in the present application has fewer model parameters, the model structure is very easy to optimize, the training and inference are very efficient, and it can be effectively extended to run in real time on low-computing-power platforms. Description of the Drawings

[0045] Figure 1Flowchart of an image de-raining method based on deformable multi-head attention in this application;

[0046] Figure 2 Schematic diagram of the network structure of the deformable Transformer model in this application;

[0047] Figure 3 Schematic diagram of the structure of the encoder in this application;

[0048] Figure 4 Schematic diagram of the structure of the basic decoder in this application;

[0049] Figure 5 Flowchart of another embodiment of an image de-raining method based on deformable multi-head attention in this application;

[0050] Figure 6 Schematic diagram of the structure of an image de-raining device based on deformable multi-head attention;

[0051] Figure 7 Schematic diagram of the hardware structure of an image de-raining device based on deformable multi-head attention. Detailed implementation manners

[0052] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts shall fall within the protection scope of this application.

[0053] To make the purpose, technical solutions and advantages of this application clearer, the embodiments of this application will be further described in detail below in conjunction with the accompanying drawings.

[0054] In a first aspect, the embodiments of this application provide an image de-raining method based on deformable multi-head attention, which uses an encoder-decoder architecture, proposes a Transformer module based on a deformable multi-head self-attention mechanism, and applies it to the encoder-decoder to construct an end-to-end deep network, efficiently realizing high-quality image de-raining in multiple scenarios.

[0055] In one embodiment, refer to Figure 1 , Figure 1 which is the schematic flowchart of the image de-raining method based on deformable multi-head attention in this application. As Figure 1 shown, the image de-raining method based on deformable multi-head attention includes:

[0056] S1: Continuously downsample the image to be processed based on a preset encoder to obtain the deep semantic features of the image to be processed;

[0057] S2: Upsample the depth semantic features based on a preset decoder to perform the rain removal operation on the image to be processed; specifically, upsample the depth semantic features of the image to be processed obtained by the decoder, restore the image to be processed to its original size, and enhance the efficient embedding of the rain-free semantic features during this process to adaptively achieve image rain removal.

[0058] It should be noted that the encoder is composed of multiple deformable Transformer models (a natural language processing model) connected in series. The deformable Transformer model is constructed based on deformable multi-head self-attention and downsampling, and the decoder is created based on convolution and transposed convolution. That is, connect the deformable Transformer models in series to form the encoder, and connect convolution and transposed convolution in series to form the decoder.

[0059] Further, in one embodiment, the encoder is composed of a set number of deformable Transformer models connected in series, and the output of each deformable Transformer model is used as the input of the next deformable Transformer model. The output of the last deformable Transformer model is used as the output of the decoder; the decoder includes a set number of basic decoders connected in series, and the basic decoder is composed of convolution and transposed convolution connected in series; the input of the first basic decoder in the decoder is the output of the last deformable Transformer model in the encoder, and the inputs of other basic decoders include a first part of the input and a second part of the input; the first part of the input is the output of the previous basic decoder, and the second part of the input is the output of the corresponding deformable Transformer model in the encoder; the output of the last basic decoder in the decoder is the rain-removed image corresponding to the image to be processed.

[0060] Specifically, for any input image, its image matrix data can be represented in the form of a tensor as Img ∈ R 3×H×W , where H is the height of the image, W is the width of the image, the encoder is composed of 4 deformable Transformer models described in this application connected in series, the output of each deformable Transformer model will be used as the input of the next deformable Transformer model, and each deformable Transformer model reduces the scale of the input feature tensor to 1 / 4 of the original and expands the channel dimension to 2 times the original. The output of the i-th deformable Transformer model is expressed as where d i = 2 i , H i = H / 2 i , W i = W / 2 i . FM4 As the input of the decoder, it is passed to the decoder for Transformer mapping, and the spatial scale of the output feature map is the same as that of FM 4 Consistent, channel dimension is FM 4 Finally, the decoder output is upsampled through multiple serial modules based on deconvolution and deformable convolution to obtain a clean (enhanced) image. i Where i=1,2,3, and is input into the basic decoder corresponding to the decoder through skip connection.

[0061] Furthermore, in one embodiment, for image patch feature generation, specifically: for any input image, its image matrix data can be represented in the form of a tensor as Img∈R 3×H×W , use the deep separable large kernel convolution proposed by ConvNeXt to extract features from the input image, use a large kernel convolution with a convolution kernel of 31 and a step size of 2 to map Img to The tensor d is the channel dimension of the output tensor. The use of large kernel convolution can capture the non-local semantic association of the image while retaining the image details as much as possible, thus improving the semantic expression ability of shallow features.

[0062] Furthermore, in one embodiment, the deformation Transformer model is composed of deformation multi-head self-attention and downsampling modules, and the deformation multi-head self-attention includes a global deformation multi-head self-attention mechanism and a local deformation multi-head self-attention mechanism. Each deformation Transformer model is connected in series with regularization and multi-layer perceptron operations. The network structure of the deformation Transformer model is as follows: Figure 2 As shown in Figure 1, four series-connected deformation transformer models form the backbone network of the encoder. In this application, a deformation transformer model is referred to as a stage.

[0063] Downsampling is implemented based on the PatchMering module (for downsampling) in the Swin Transformer (a hierarchical structure). Therefore, for any input tensor In∈R d×H×W , after Transformer Block mapping, the dimension of the output tensor becomes The overall structure of the encoder is as follows Figure 3 shown.

[0064] Furthermore, in one embodiment, for the global deformation multi-head self-attention mechanism, the specific implementation method is:

[0065] S101: Based on the deformation sampling mechanism in the deformable convolution algorithm, construct a feature graph deformation transformation model; specifically, for any input tensor In ∈ R d×H×W , first use a deformation sampling mechanism similar to that in the deformable convolution algorithm to construct a feature graph deformation transformation model.

[0066] S102: Process the input tensor through multiple cascaded feature graph deformation transformation models, so that the elements in the input tensor are redistributed according to semantic information to obtain the processed input tensor;

[0067] Specifically, after In is processed by two cascaded feature graph deformation transformation models, the elements in the In tensor have been redistributed according to their own semantic information and become In d ∈R d×H×W .

[0068] S103: Average the processed input tensor in the spatial dimension into multiple windows of a set size, and use the standard multi-head attention mechanism model to process the processed input tensor in each window to obtain the output of the global deformation multi-head self-attention mechanism.

[0069] Specifically, the tensor In d is averaged into windows of size h×w in the spatial dimension, where h is the unit height and w is the unit width. The standard multi-head attention mechanism model is used to process In d in each window to obtain the output Out g of the global deformation multi-head self-attention mechanism, whose dimension is d 0 ×H×W.

[0070] Furthermore, in one embodiment, for the local deformation multi-head self-attention mechanism, the specific implementation method is:

[0071] S201: Divide the input tensor into multiple windows of a set size in the spatial dimension, and process each window with a feature graph deformation transformation model to obtain multiple deformed window regions;

[0072] Specifically, for any input tensor In ∈ R d×H×W , first average the tensor In into windows of size h×w in the spatial dimension, and process each window with a feature graph deformation transformation model to obtain deformed window regions.

[0073] S202: Execute the standard multi-head attention mechanism model in each deformed window region to obtain multiple tensors;

[0074] Specifically, the standard multi-head attention mechanism model is executed in each deformation window area to obtain tensors.

[0075] S203: Concatenate the obtained multiple tensors in the spatial dimension to obtain an output tensor with a set dimension. Specifically, tensors are concatenated in the spatial dimension, and finally an output tensor with a dimension of d×H×W is formed.

[0076] It can be seen that the four deformation Transformer models of the encoder take the input tensor In∈R d×H×W , and the output is 1 output tensor and 3 intermediate feature tensors E m1 , E m2 , E m3 , and their tensor dimensions are respectively E out input to the decoder, E m1 , E m2 , E m3 are respectively input to the corresponding basic decoder.

[0077] Furthermore, in an embodiment, the specific calculation method of the standard multi-head attention mechanism model is:

[0078]

[0079] where MSA(x) represents the output of the standard multi-head attention mechanism model, softmax represents the normalized exponential function, Q(x), k(x), V(x) represent linear transformations, x represents the input feature map, T represents matrix transpose, and d represents the channel dimension.

[0080] Since the spatial dimensions of the input and output tensors of the deformable multi-head self-attention do not change, the output tensor dimensions of the deformable multi-head self-attention within each window are all h×w, and the channel dimension is d. Thus, output tensors of d×h×w are obtained, and then by using the dimension transformation operation, an output tensor with a dimension of d×h×w can be obtained.

[0081] Furthermore, in an embodiment, the feature graph deformation transformation model is used to convert any input tensor into an output tensor with the same dimension; specifically, for any input tensor En∈R d×H×W , the feature graph deformation transformation model converts En into an output tensor On∈R d×H×W .

[0082] For converting into an output tensor with the same dimension, the specific steps include:

[0083] S301: Filter the input tensor using a convolution with a set step size, convolution kernel, and output dimension to calculate the bias tensor;

[0084] That is, first use a convolution with a step size of 1, a convolution kernel of k×k, and an output dimension of 2 to filter En, and calculate the bias tensor Pa ∈ R 2×H×W .

[0085] S302: For any position in the input tensor, determine the corresponding position in the output tensor to obtain the output tensor. Among them, for any position in the input tensor and the corresponding position in the output tensor, the pixel relationship between the two positions is:

[0086] On(p 0 ) = ∑En(p 0 + Δp 0 )

[0087] Where p 0 represents any position in the input tensor En, On(p 0 ) represents the position in the output tensor On corresponding to the position p 0 , Δp 0 represents the bias value at position p 0 , corresponding to the value at position p 0 in the bias tensor.

[0088] For any position p 0 in En, there is a corresponding position in On, and the pixel relationship between the two is as shown above.

[0089] Δp 0 is a two-dimensional vector, representing the bias situation of p 0 in the x and y axis directions respectively, thus corresponding to the value at position p 0 in the Pa tensor. Since Δp 0 is very likely to be a decimal, the present application uses linear interpolation to calculate En(p 0 ), and the interpolation formula is specifically:

[0090]

[0091] Where p represents any decimal position, q is the index of all integer positions in En, En(p) is the eigenvalue (tensor value) at position q, and G represents the bilinear interpolation function. Through the above steps, the output tensor On ∈ R d×H×W can be calculated.

[0092] Furthermore, in one embodiment, the present application forms a basic decoder (abbreviated as block) in the form of a series connection of convolution + transposed convolution. The entire decoder is composed of 4 basic decoders connected in series, and the structure of the basic decoder is as shown in Figure 4 shown. The decoder is used to restore the output features of the encoder to the input image scale.

[0093] The result output by the last basic decoder is the enhanced image, which has the same resolution as the input image. Due to the excellent generalization ability and outstanding image feature expression ability of the encoder, even using a common transposed convolution module as the decoder can achieve excellent image enhancement results.

[0094] Based on the encoder-decoder structure, the present application constructs an image enhancement model, proposes a new Deformable Transformer feature learning structure, and uses it as the encoder to continuously downsample the image to learn the deep semantic features of the image; uses a decoder based on transposed convolution to upsample the learned features and restore the image to the original size. During this process, the efficient embedding of rain-free semantic features is enhanced to adaptively achieve image deraining. In Deformable Transformer, the present application improves the deformation sampling strategy within the global and local receptive fields and efficiently implements the deformable multi-head self-attention mechanism in the form of divided deformation windows. The present application has the following advantages:

[0095] (1) It is applicable to image deraining in a variety of rainy weather and scenarios, and has a low computational complexity. The existing methods generally face two bottlenecks. One is that the deraining scenarios of the model are limited. Most models are difficult to have the same excellent deraining effect on rain lines and raindrops attached to the screen, and there are obvious differences in the deraining effect under different lighting scenarios. The other is that the model complexity is relatively high. These methods mostly rely on stacking deep models to enhance the deraining effect and generalization ability of the model, consuming a high amount of computing resources, which affects the practicality and real-time performance of the technology. The proposed encoder-decoder structure in the present application has a relatively simple model structure and limited computing power consumption. At the same time, the proposed deformable multi-head attention mechanism in the present application can effectively improve the feature expression ability of the encoder with a very limited additional computing power consumption, thereby improving the deraining effect of the overall model.

[0096] (2) The encoder can focus more on the feature learning of the non-rain area, and reduce the influence of raindrop / rain line noise on the deraining result in a simple but efficient way, thereby achieving high-quality image deraining for multiple scenarios and multiple raindrop types.

[0097] (3) The encoder deformable multi-head self-attention mechanism can capture the correlation relationships between pixels with different ranges, can extract deeper features with more levels and richer semantics, and can further improve the image deraining effect.

[0098] Second aspect, refer to Figure 5 As shown, the embodiment of the present application also provides an image de-raining method based on deformable multi-head attention, which specifically includes the following steps:

[0099] a: Construct a deformable Transformer model according to deformable multi-head self-attention and downsampling, and concatenate multiple deformable Transformer models to obtain an encoder, and create a decoder based on convolution and transposed convolution;

[0100] b: Continuously downsample the image to be processed based on the encoder to obtain the deep semantic features of the image to be processed;

[0101] c: Upsample the deep semantic features based on the decoder to implement the de-raining operation on the image to be processed.

[0102] Third aspect, the embodiment of the present application also provides an image de-raining device based on deformable multi-head attention.

[0103] In one embodiment, refer to Figure 6 , Figure 6 is a schematic diagram of the functional modules of the image de-raining device based on deformable multi-head attention of the present application. As Figure 6 shown, the image de-raining device based on deformable multi-head attention includes an acquisition module and an execution module.

[0104] The acquisition module is used to continuously downsample the image to be processed based on a preset encoder to obtain the deep semantic features of the image to be processed; the execution module is used to upsample the deep semantic features based on a preset decoder to implement the de-raining operation on the image to be processed; wherein, the encoder is composed of multiple deformable Transformer models connected in series, the deformable Transformer model is constructed according to deformable multi-head self-attention and downsampling, and the decoder is created based on convolution and transposed convolution.

[0105] Among them, the functional implementation of each module in the above image de-raining device based on deformable multi-head attention corresponds to each step in the above image de-raining method embodiment based on deformable multi-head attention, and its functions and implementation processes will not be elaborated here one by one.

[0106] Fourth aspect, the embodiment of the present application provides an image de-raining device based on deformable multi-head attention. The image de-raining device based on deformable multi-head attention can be a device with data processing functions such as a personal computer (PC), a laptop computer, a server, etc.

[0107] Refer to Figure 7 , Figure 7This is a schematic diagram of the hardware structure of the image de-raining device based on deformable multi-head attention involved in the solution of the embodiment of the present application. In the embodiment of the present application, the image de-raining device based on deformable multi-head attention may include a processor, a memory, a communication interface, and a communication bus.

[0108] Among them, the communication bus can be of any type and is used to interconnect the processor, the memory, and the communication interface.

[0109] The communication interface includes input / output (I / O) interfaces, physical interfaces, and logical interfaces, etc., which are used to implement the interconnection of components inside the image de-raining device based on deformable multi-head attention, as well as interfaces used to interconnect the image de-raining device based on deformable multi-head attention with other devices (such as other computing devices or user devices). The physical interface can be an Ethernet interface, a fiber optic interface, an ATM interface, etc.; the user device can be a display, a keyboard, etc.

[0110] The memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical memory, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.

[0111] The processor can be a general-purpose processor, and the general-purpose processor can call the image de-raining program stored in the memory and execute the image de-raining method provided by the embodiment of the present application. For example, the general-purpose processor can be a central processing unit (CPU). Among them, the method executed when the image de-raining program based on deformable multi-head attention is called can refer to each embodiment of the image de-raining method based on deformable multi-head attention of the present application, and will not be elaborated here.

[0112] Those skilled in the art can understand that Figure 6 the hardware structure shown in does not constitute a limitation to the present application, and may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.

[0113] In a fifth aspect, the embodiment of the present application further provides a computer-readable storage medium.

[0114] A program for image de-raining based on deformable multi-head attention is stored on a computer-readable storage medium of the present application. When the program for image de-raining based on deformable multi-head attention is executed by a processor, the steps of the method for image de-raining based on deformable multi-head attention as described above are implemented.

[0115] Among them, the method implemented when the program for image de-raining based on deformable multi-head attention is executed can refer to the various embodiments of the method for image de-raining based on deformable multi-head attention of the present application, which will not be elaborated here.

[0116] It should be noted that the serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.

[0117] The terms "including" and "having" and any variations thereof in the description of the specification, claims and drawings of the present application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products or devices. The descriptions with terms such as "first", "second" and "third" are used to distinguish different objects, etc., and do not represent a sequence, nor do they limit that "first", "second" and "third" are different types.

[0118] In the description of the embodiments of the present application, terms such as "exemplary", "for example" or "for instance" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary", "for example" or "for instance" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of terms such as "exemplary", "for example" or "for instance" is intended to present relevant concepts in a specific manner.

[0119] In the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may represent A or B; the "and / or" in the text is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present application, "a plurality of" means two or more than two.

[0120] In some processes described in the embodiments of the present application, there are multiple operations or steps that appear in a specific order. However, it should be understood that these operations or steps may not be executed in the order in which they appear in the embodiments of the present application or may be executed in parallel. The serial numbers of the operations are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations or steps may be executed in sequence or in parallel, and these operations or steps may be combined.

[0121] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above and includes several instructions for causing a terminal device to execute the methods described in the various embodiments of the present application.

[0122] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. An image de-raining method based on deformable multi-head attention, characterized in that, the image de-raining method based on deformable multi-head attention includes: Performing continuous downsampling processing on the image to be processed based on a preset encoder to obtain the deep semantic features of the image to be processed; Performing upsampling processing on the deep semantic features based on a preset decoder to implement the de-raining operation on the image to be processed; Among them, the encoder is composed of multiple deformable Transformer models connected in series, the deformable Transformer model is constructed according to deformable multi-head self-attention and downsampling, and the decoder is created based on convolution and transposed convolution; Among them, the deformable multi-head self-attention includes a global deformable multi-head self-attention mechanism and a local deformable multi-head self-attention mechanism; the downsampling is implemented based on the PatchMering module in Swin Transformer; Among them, for the global deformable multi-head self-attention mechanism, the specific implementation method is: Processing the input tensor through multiple serially connected feature map deformation transformation models, so that the elements in the input tensor are redistributed according to semantic information to obtain the processed input tensor; The processed input tensor is evenly divided into multiple windows of a set size in the spatial dimension, and the standard multi-head attention mechanism model is used to process the processed input tensor in each window to obtain the output of the global deformable multi-head self-attention mechanism; Among them, the feature map deformation transformation model is constructed based on the deformation sampling mechanism in the deformable convolution algorithm.

2. The image de-raining method based on deformable multi-head attention according to claim 1, characterized in that, for the local deformable multi-head self-attention mechanism, the specific implementation method is: Dividing the input tensor into multiple windows of a set size in the spatial dimension, and processing each window with a feature map deformation transformation model to obtain multiple deformed window regions; Executing the standard multi-head attention mechanism model in each deformed window region to obtain multiple tensors; The multiple obtained tensors are concatenated in the spatial dimension to obtain an output tensor of a set dimension.

3. The image de-raining method based on deformable multi-head attention according to claim 2, characterized in that: The feature map deformation transformation model is used to convert any input tensor into an output tensor of the same dimension; For the output tensor converted into the same dimension, the specific steps include: Filtering the input tensor with a convolution with a set stride, convolution kernel and output dimension to calculate the bias tensor; For any position in the input tensor, determining the corresponding position in the output tensor, so as to obtain the output tensor.

4. The image de-raining method based on deformable multi-head attention according to claim 1, characterized in that: In the encoder, the output of each deformable Transformer model is used as the input of the next deformable Transformer model, and the output of the last deformable Transformer model is used as the output of the decoder; The decoder includes a set number of basic decoders connected in series, and the basic decoder is composed of convolution and transposed convolution connected in series; The input of the first basic decoder in the decoder is the output of the last deformable Transformer model in the encoder, and the inputs of other basic decoders include a first part of the input and a second part of the input; The first part of the input is the output of the previous basic decoder, and the second part of the input is the output of the corresponding deformable Transformer model in the encoder; The output of the last basic decoder in the decoder is the de-rained image corresponding to the image to be processed.

5. An image de-raining method based on deformable multi-head attention, characterized in that, the image de-raining method based on deformable multi-head attention includes: Constructing a deformable Transformer model according to deformable multi-head self-attention and downsampling, and concatenating multiple deformable Transformer models to obtain an encoder, and creating a decoder based on convolution and transposed convolution; Performing continuous downsampling processing on the image to be processed based on the encoder to obtain the deep semantic features of the image to be processed; Performing upsampling processing on the deep semantic features based on the decoder to implement the de-raining operation on the image to be processed; wherein, the deformable multi-head self-attention includes a global deformable multi-head self-attention mechanism and a local deformable multi-head self-attention mechanism; the downsampling is implemented based on the PatchMering module in Swin Transformer; wherein, for the global deformable multi-head self-attention mechanism, the specific implementation method is: Processing the input tensor through multiple cascaded feature map deformation transformation models so that the elements in the input tensor are redistributed according to semantic information to obtain the processed input tensor; Dividing the processed input tensor into multiple windows of a set size on average in the spatial dimension, and processing the processed input tensor in each window using a standard multi-head attention mechanism model to obtain the output of the global deformable multi-head self-attention mechanism; wherein, the feature map deformation transformation model is constructed based on the deformation sampling mechanism in the deformable convolution algorithm.

6. An image de-raining device based on deformable multi-head attention, characterized in that, the image de-raining device based on deformable multi-head attention includes: An acquisition module, which is used to perform continuous downsampling processing on the image to be processed based on a preset encoder to obtain the deep semantic features of the image to be processed; An execution module, which is used to perform upsampling processing on the deep semantic features based on a preset decoder to implement the de-raining operation on the image to be processed; wherein, the encoder is composed of multiple deformable Transformer models connected in series, the deformable Transformer model is constructed according to deformable multi-head self-attention and downsampling, and the decoder is created based on convolution and transposed convolution; wherein, the deformable multi-head self-attention includes a global deformable multi-head self-attention mechanism and a local deformable multi-head self-attention mechanism; the downsampling is implemented based on the PatchMering module in Swin Transformer; wherein, for the global deformable multi-head self-attention mechanism, the specific implementation method is: Process the input tensor through multiple cascaded feature graph transformation models so that the elements in the input tensor are redistributed according to semantic information, and obtain the processed input tensor; Average the processed input tensor in the spatial dimension into multiple windows of a set size, and use the standard multi-head attention mechanism model to process the processed input tensor in each window to obtain the output of the global deformation multi-head self-attention mechanism; Among them, the feature graph transformation model is constructed based on the deformation sampling mechanism in the deformation convolution algorithm.

7. An image de-raining device based on deformed multi-head attention, characterized in that, The image de-raining device based on deformed multi-head attention includes a processor, a memory, and an image de-raining program based on deformed multi-head attention stored on the memory and executable by the processor. When the image de-raining program based on deformed multi-head attention is executed by the processor, the steps of the image de-raining method based on deformed multi-head attention as described in any one of claims 1 to 4 are implemented.

8. A computer-readable storage medium, characterized in that, An image de-raining program based on deformed multi-head attention is stored on the computer-readable storage medium. When the image de-raining program based on deformed multi-head attention is executed by a processor, the steps of the image de-raining method based on deformed multi-head attention as described in any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Single image rain removal method based on multi-stage feature complementation network

    CN113962905A

  • Image rain removal method and device

    CN117078574A

  • Image rain removal method based on multi-scale and multi-head attention

    CN117151999A