Method and device for removing rain from UAV video images based on latent frequency domain representation
Through the method based on potential frequency domain characterization, the contradiction between rain pattern elimination and detail retention in drone video image rain removal technology is solved, and efficient image clarity is improved in rainy environments, ensuring the rationality of image contrast and visual structure.
Patent Information
- Application Number
- CN202510252135.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-05
AI Technical Summary
When facing actual scenes with uncertain rain density distribution, existing drone video image rain removal technology is difficult to balance the contradiction between rain pattern elimination and detail retention, resulting in high-frequency information loss and texture distortion, and poor rain removal effect.
Using a method based on latent frequency domain characterization, the extraction network and dual-frequency potential characterization enhancement through image comparison regularization processing, latent frequency domain characterization extraction network and dual-frequency potential characterization enhancement are identified, high-frequency components related to rain fringe noise in the image, low-frequency components related to the overall structure and background of the image are retained, image details and texture information are enhanced, and the rain removal effect of the image is finally achieved.
It improves the clarity of drone video images on rainy days, effectively removes raindrop occlusion and background blur, maintains reasonable contrast and visual structure of the image, and improves the visual effect of the image.
Smart Images

Figure CN119850484B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to a method and device for removing rain from UAV video images based on latent frequency domain representation. Background Art
[0002] The technology of removing rain from UAV video images has important applications in multiple fields, such as video surveillance, autonomous driving, medical imaging, etc. It can improve the clarity of images and ensure the normal operation of the vision system. However, the video images captured by UAVs in bad weather conditions (such as rainy days) will be interfered by raindrops, resulting in blurred images and reduced contrast, seriously affecting the performance of the vision system.
[0003] Current research on single-image rain removal focuses on solving the interference of dense rain streaks and their stacking effects on the restoration of background details, and retaining image content while removing rain streaks through a progressive elimination strategy. However, existing methods are generally limited by the learning paradigm of fixed rain streak features. When facing the actual scenario with uncertain rain density distribution, a unified feature extraction mechanism is usually adopted to achieve model generality, rather than dynamically adjusting the processing strategy according to the rain streak density. This one-size-fits-all method is prone to over-rain-removal phenomena in high-density rain areas. Especially when there are fine-grained objects in the background, it is difficult to balance the contradiction between rain streak elimination and detail retention, resulting in problems such as loss of high-frequency information and texture distortion. Therefore, how to improve the effect of removing rain from UAV video images has become an urgent problem to be solved. Summary of the Invention
[0004] The present invention provides a method and device for removing rain from UAV video images based on latent frequency domain representation, and its main purpose is to solve the problem of poor effect of removing rain from UAV video images.
[0005] To achieve the above object, a method for removing rain from UAV video images based on latent frequency domain representation provided by the present invention includes: acquiring complex dynamic video images of a UAV in rainy weather, performing image contrast regularization processing on the complex dynamic video images of the UAV in rainy weather to obtain a contrast-regularized image; guiding a pre-constructed latent frequency domain representation extraction network according to the contrast-regularized image to extract the latent frequency domain representation of the complex dynamic video images of the UAV in rainy weather; performing dual-frequency latent representation enhancement on the latent frequency domain representation to obtain an enhanced representation; and performing representation aggregation image restoration according to the enhanced representation to obtain a rain-removed image of the complex dynamic video images of the UAV in rainy weather.
[0006] The present invention also provides a rain removal device for UAV video images based on latent frequency domain representation, including: an image contrast regularization processing module, configured to obtain complex dynamic video images of a UAV in rainy days, perform image contrast regularization processing on the complex dynamic video images of the UAV in rainy days to obtain a contrast-regularized image; a latent frequency domain representation extraction module, configured to extract the latent frequency domain representation of the complex dynamic video images of the UAV in rainy days according to the contrast-regularized image to guide a pre-constructed latent frequency domain representation extraction network; a dual-frequency latent representation enhancement module, configured to perform dual-frequency latent representation enhancement on the latent frequency domain representation to obtain an enhanced representation; and an image restoration module, configured to perform representation aggregation image restoration according to the enhanced representation to obtain a rain-removed image of the complex dynamic video images of the UAV in rainy days.
[0007] In an embodiment of the present invention, by performing image contrast regularization processing on the complex dynamic video images of the UAV in rainy days to obtain a contrast-regularized image, the texture information and detail information of the image can be highlighted, and the reasonable contrast and visual structure of the image can be ensured during the processing; according to the contrast-regularized image to guide a pre-constructed latent frequency domain representation extraction network to extract the latent frequency domain representation of the complex dynamic video images of the UAV in rainy days, the high-frequency components related to rain stripe noise in the image can be identified, and the low-frequency components related to the overall structure and background of the image can be retained, effectively extracting a robust latent frequency domain representation; performing dual-frequency latent representation enhancement on the latent frequency domain representation to obtain an enhanced representation can enhance the latent frequency domain representation from both low-frequency and high-frequency aspects, thereby effectively focusing on the low-frequency local texture details and object contours in the complex dynamic images of the UAV in rainy days and improving the effect of subsequent image restoration; performing representation aggregation image restoration according to the enhanced representation to achieve a robust and efficient video image rain removal effect in the complex dynamic scenario of the UAV in rainy days, ultimately improving the visual effect of the image and obtaining a more accurate rain-removed image. Therefore, a method and device for rain removal of UAV video images based on latent frequency domain representation proposed by the present invention can solve the problem of poor rain removal effect of UAV video images. Description of the Drawings
[0008] Figure 1 It is a schematic flowchart of a method for rain removal of UAV video images based on latent frequency domain representation provided by an embodiment of the present invention;
[0009] Figure 2 It is a functional module diagram of a residual wavelet transform convolution provided by an embodiment of the present invention;
[0010] Figure 3 It is a functional module diagram of an image contrast regularization prior module provided by an embodiment of the present invention;
[0011] Figure 4 It is a functional module diagram of a frequency domain state space module provided by an embodiment of the present invention;
[0012] Figure 5 The functional module diagram of the frequency domain characterization mining module provided by an embodiment of the present invention;
[0013] Figure 6 The functional module diagram of the contrast prior refinement module provided by an embodiment of the present invention;
[0014] Figure 7 The functional module diagram of the prior gating feedforward module provided by an embodiment of the present invention;
[0015] Figure 8 The functional module diagram of the depth frequency domain characterization aggregation image restoration network provided by an embodiment of the present invention;
[0016] Figure 9 The network architecture diagram of the rain removal framework for complex dynamic video images of drones in rainy days based on latent frequency domain characterization provided by an embodiment of the present invention;
[0017] Figure 10 The functional module diagram of the device for removing rain from drone video images based on latent frequency domain characterization provided by an embodiment of the invention.
[0018] The implementation, functional features, and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. Specific Embodiments
[0019] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0020] An embodiment of the present application provides a method for removing rain from drone video images based on latent frequency domain characterization. The execution subject of the method for removing rain from drone video images based on latent frequency domain characterization includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the method for removing rain from drone video images based on latent frequency domain characterization can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0021] Refer to Figure 1 As shown, it is a flowchart of a method for removing rain from drone video images based on latent frequency domain characterization provided by an embodiment of the present invention.
[0022] S1. Obtain the complex dynamic video images of the drone in rainy days, perform image contrast regularization processing on the complex dynamic video images of the drone in rainy days, and obtain the contrast-regularized images.
[0023] In the embodiments of the present invention, the complex dynamic video images of the drone in rainy days are the dynamic video images captured by the drone in rainy days. However, the rainy-day images collected by the drone are often severely blocked by rainwater and the background is blurred. Therefore, image processing is required to obtain the images after removing the rain.
[0024] Specifically, performing image contrast regularization processing on the complex dynamic video images of the drone in rainy days to obtain contrast-regularized images includes: using a pre-constructed image contrast regularization prior module to extract the contrast of the complex dynamic video images of the drone in rainy days; performing contrast regularization operation on the complex dynamic video images of the drone in rainy days according to the contrast to obtain the regularized contrast; performing residual wavelet transform convolution on the regularized contrast to obtain the contrast-regularized images.
[0025] Specifically, the image contrast regularization prior module (Image Contrast - Regularization Prior Module, ICPM) calculates the contrast by fully extracting the maximum and minimum channel values in the complex dynamic video images of the drone in rainy days, and highlights the image texture information and detail information through regularization operation and convolution operation, so as to provide image contrast prior information for the subsequent potential frequency domain characterization extraction network, ensure that the image maintains reasonable contrast and visual structure during the processing, and avoid over-smoothing or information loss.
[0026] Specifically, the image contrast regularization prior module can be expressed as:
[0027]
[0028]
[0029]
[0030]
[0031] Among them, represents the complex dynamic video images of the drone in rainy days, represents function, represents the minimum value of the image channels of the complex dynamic video images of the drone in rainy days, represents the maximum value of the image channels of the complex dynamic video images of the drone in rainy days, represents the contrast regularization operation, represents 1×1 convolution, represents the residual wavelet transform convolution, Represents the contrast regularized image. It is worth mentioning that the undefined parameters in the above formula are all intermediate parameters used in the calculation process.
[0032] Furthermore, UAV platforms often have characteristics such as strong mobility, which makes the real-time performance requirements of algorithm deployment strict. In this regard, the method of the present invention has fully considered the unique characteristics of the deployment platform at the beginning of the design of the neural network model architecture, taking the lightweight design of the model as an important starting point, and reducing the number of model parameters as the main means to reduce the computational load at runtime as much as possible to adapt to the limited computing power resources of the UAV platform. In order to reduce the number of model parameters without affecting the performance of the model as much as possible, improve its operating efficiency, and thus adapt to the scenario where the computing resources of the UAV platform are limited, the method of the present invention adopts a residual wavelet transform convolution method, namely RWTConv, such as Figure 2 shown.
[0033] Figure 2 In the convolution method, WTConv (Wavelet Transform Convolution) is a convolution method that combines wavelet transform and traditional convolution operations. It can expand the receptive field of convolution and improve feature extraction capabilities. WTConv uses wavelet transform to perform multi-level decomposition of input data, decomposing the image into different frequency components (such as low frequency and high frequency), and then performs small-size convolution operations on each frequency layer. Finally, the results are recombined through inverse wavelet transform. In this way, it can effectively capture low-frequency information in the image and expand the receptive field of convolution without significantly increasing the number of parameters; Conv1x1 is a one-dimensional convolution. represents an element-wise multiplication operation, represents the tensor addition operation, and GhostConv represents the ghost convolution module.
[0034] In detail, the Residual Wavelet Transform Convolution (RWTConv) module is used as the basic operation for building the model. This operation uses wavelet transform to process local information in the frequency domain, and can provide a larger receptive field with the same number of parameters. In addition, it combines 1×1 convolution and identity mapping in a convolution layer to maintain the uniformity of parameters, and uses the residual structure to reduce the loss of low-dimensional potential frequency domain representation information. Finally, the potential frequency domain representation information is integrated through the convolution operation. This operation has great advantages in reducing the number of model parameters while maintaining model performance. Its specific operations are as follows: Figure 2 As shown, for any input tensor , which is mapped to the output tensor through residual wavelet transform convolution .
[0035] In the embodiments of the present invention, the residual wavelet transform convolution is expressed as:
[0036]
[0037] Wherein, represents the input tensor of the residual wavelet transform convolution, represents the wavelet transform convolution, is a 1×1 convolution, represents the phantom convolution, represents the element-wise multiplication operation, represents the tensor addition operation, represents the output tensor of the residual wavelet transform convolution. It is worth mentioning that the parameters not explained in the above formula are intermediate parameters used in the calculation process.
[0038] Specifically, the image contrast regularization prior module can be as Figure 3 shown, wherein, represents the maximum value of the image channels, represents the minimum value of the image channels, represents the Sigmoid function, represents the contrast regularization operation, represents a 1×1 convolution, represents the residual wavelet transform convolution, in the figure represents the contrast regularization operation, represents the element-wise multiplication operation, represents the tensor addition operation.
[0039] In the embodiments of the present invention, by performing image contrast regularization processing on the complex dynamic video images of the unmanned aerial vehicle in rainy days, the image texture information and detail information can be highlighted, and it is ensured that the image maintains a reasonable contrast and visual structure during the processing, improving the effect of image de-raining.
[0040] S2. Extract the latent frequency domain representation of the complex dynamic video image of the unmanned aerial vehicle in rainy days according to the latent frequency domain representation extraction network pre-constructed by the contrast-regularized image guidance.
[0041] In the embodiments of the present invention, aiming at the problems of generally low object resolution, serious rain streak occlusion, and background blurring in the complex dynamic images of unmanned aerial vehicles (UAVs) on rainy days, the present invention designs a contrast-aware frequency-domain latent-representation extraction network (CFRNet), which converts image information into the frequency domain, identifies the high-frequency components related to rain streak noise in the image, and retains the low-frequency components related to the overall structure and background of the image, so as to effectively extract robust latent frequency-domain representations. In addition, the network makes full use of the prior knowledge of image contrast to avoid the loss of key original image information and enhance the image detail and texture information in the robust latent frequency-domain representation. The latent frequency-domain representation extraction network consists of a frequency-domain state space module (FSSB) and multiple repeating units, and each unit is composed of a frequency-domain representation mining module (FRMM), a prior-gated feed forward module (PFFM), and a contrast prior refinement module (CPRM).
[0042] Specifically, extracting the latent frequency-domain representation of the complex dynamic video image of the UAV on a rainy day according to the latent frequency-domain representation extraction network pre-constructed by contrast-regularized image guidance includes: using the frequency-domain state space module in the latent frequency-domain representation extraction network to extract the image frequency-domain representation in the complex dynamic video image of the UAV on a rainy day; using the frequency-domain representation mining module in the latent frequency-domain representation extraction network to calculate the spectral robust latent representation of the complex dynamic video image of the UAV on a rainy day; using the contrast prior refinement module in the latent frequency-domain representation extraction network to perform representation mining on the contrast-regularized image to obtain the mined representation; using the prior-gated feed forward module in the latent frequency-domain representation extraction network to perform gated learning on the spectral robust latent representation and the mined representation to obtain the gated representation; and generating the latent frequency-domain representation of the complex dynamic video image of the UAV on a rainy day according to the image frequency-domain representation and the gated representation.
[0043] Specifically, to enhance the potential frequency-domain representation learning ability of UAV rainy-day complex dynamic images, the present invention proposes a frequency-domain state space module FSSB. Among them, FSSB is mainly composed of a linear embedding layer, a 2D selective scanning module SS2D, and a potential frequency-domain representation learning layer. The frequency-domain state space module can effectively traverse the frequency-domain information of the image, thereby modeling a robust potential frequency-domain representation. In addition, the frequency-domain state space module can achieve linear complexity without sacrificing the global receptive field, improve the computational efficiency, and facilitate the edge deployment of the inventive method on a UAV platform with limited computational resources and payload. The specific structure is as shown in Figure 4 shown. Among them, LayerNorm(·) represents the layer normalization operation, Linear(·) represents the linear mapping, FFT2D(·) represents the two-dimensional fast Fourier transform, RWTConv(·) represents the residual wavelet transform convolution, GeLU(·) represents the GeLU activation function, InvFFT2D(·) represents the two-dimensional inverse fast Fourier transform, SS2D(·) represents the 2D selective scanning module, represents the element-wise multiplication operation, represents the tensor addition operation.
[0044] Specifically, the frequency-domain state space module can be expressed as:
[0045]
[0046] Among them, represents the UAV rainy-day complex dynamic video image, represents the layer normalization operation, represents the linear mapping, represents the two-dimensional fast Fourier transform, represents the residual wavelet transform convolution, represents the GeLU activation function, represents the two-dimensional inverse fast Fourier transform, represents the 2D selective scanning module, represents the image frequency-domain representation. It should be noted that the parameters not explained in the above formula are all intermediate parameters used in the calculation process.
[0047] Furthermore, to effectively focus on different rain trace characterizations in the complex dynamic video images of drones in rainy days and conduct robust latent characterization modeling in the image spectrum, the method of the present invention proposes a frequency-domain characterization mining module FRMM for mining multi-receptive field information of the latent frequency-domain characterization of the input regularized image. Specifically, the frequency-domain characterization mining module first uses residual wavelet transform convolution for preliminary characterization modeling, then pays attention to rain noises in different directions through asymmetric convolution, improves the characterization modeling performance while achieving lightweight, and then effectively extracts the image frequency-domain information through operations such as Fourier transform, making the modeled latent frequency-domain characterization more conducive to the subsequent rain removal process. The specific structure is as Figure 5 shown, where represents residual wavelet transform convolution, represents the separation operation, represents asymmetric convolution, represents the ReLU activation function, represents the connection operation in the channel dimension, represents 1×1 convolution, represents the two-dimensional fast Fourier transform, represents the GeLU activation function, represents the inverse two-dimensional fast Fourier transform. In the figure, represents the connection in the channel dimension and the convolution dimensionality reduction operation, represents the tensor addition operation.
[0048] Specifically, the frequency-domain characterization mining module can be expressed as:
[0049]
[0050] where represents the complex dynamic video image of a drone in rainy days, represents residual wavelet transform convolution, represents the separation operation, represents asymmetric convolution, represents the ReLU activation function, represents the connection operation in the channel dimension, represents 1×1 convolution, represents the two-dimensional fast Fourier transform, represents the GeLU activation function, represents the inverse two-dimensional fast Fourier transform, represents the connection in the channel dimension and the convolution dimensionality reduction operation, represents the spectral robust latent characterization. It should be noted that the parameters not explained in the above formula are all intermediate parameters used in the calculation process.
[0051] Furthermore, in order to continuously enrich the channel dimension information of the contrast of complex dynamic video images of drones in rainy days, so as to continuously supplement low-frequency details and local textures from shallow to deep in the process of iteratively extracting the potential frequency-domain representation of the images, avoid over-smoothing or information loss, and effectively retain the key original clear structure of the images, the present invention proposes a contrast prior refinement module CPRM. The contrast prior refinement module enhances local contrast information through residual wavelet transform convolution and retains the original image contrast information through a residual structure to prevent potential interference caused by an overly deep network. Finally, the integrated image contrast prior knowledge is transmitted to the corresponding prior gating feed-forward module.
[0052] Specifically, the contrast prior refinement module can be expressed as:
[0053]
[0054] Among them, represents the contrast-regularized image, represents the residual wavelet transform convolution, represents the asymmetric convolution, represents function, represents the 1×1 convolution, represents the soft pooling operation, represents the mined representation. It should be noted that the parameters not interpreted in the above formula are all intermediate parameters used in the calculation process.
[0055] Among them, for the specific structure of the contrast prior refinement module, please refer to Figure 6 .
[0056] Furthermore, to effectively utilize the contrast prior knowledge of complex dynamic images of drones in rainy days, the method of the present invention designs a prior gating feed-forward module PFFM. The prior gating feed-forward module integrates the contrast prior knowledge of the image into the modeled potential frequency-domain representation in a gating manner to improve local details and structure recovery. Specifically, the prior gating feed-forward module first enriches the channel dimension through residual wavelet transform convolution, then uses asymmetric convolution to refine local details and object contours. At the same time, the prior branch introduces the enhanced image contrast information and uses this information as the gating weight, thereby significantly strengthening and reconstructing the image background information weakened by rain noise. The specific structure is as Figure 7 shown, where represents the residual wavelet transform convolution, represents the asymmetric convolution, represents the Tanh activation function, represents the 1×1 convolution, represents the gating unit, Figure 7 in represents the gating unit, Denotes an element-wise multiplication operation, Denotes a tensor addition operation.
[0057] The prior gating feed-forward module can be expressed by the formula:
[0058]
[0059] Where, Denotes the spectral robust latent representation, Denotes the mined representation, Denotes the residual wavelet transform convolution, Denotes the asymmetric convolution, Denotes the Tanh activation function, Denotes the 1×1 convolution, Denotes the gating unit, Denotes the gated representation. It is worth mentioning that the parameters not defined in the above formula are all intermediate parameters used in the calculation process.
[0060] In the embodiments of the present invention, the number of units N of the frequency-domain representation mining module, the prior gating feed-forward module, and the contrast prior refinement module can be adaptively adjusted according to the rain streak density and raindrop occlusion degree of the specific task environment of the rainy-day complex dynamic video image of the UAV, so as to extract the robust latent frequency-domain representation from the rainy-day complex dynamic video image of the UAV for subsequent representation enhancement.
[0061] Each unit can be expressed as:
[0062]
[0063] Where, Is the frequency-domain representation mining module, Is the contrast prior refinement module, Is the contrast prior refinement module.
[0064] In the embodiments of the present invention, the image frequency-domain representation and the gated representation are subjected to element-wise multiplication processing to obtain the latent representation in the frequency domain, that is, the latent frequency-domain representation. Through the latent frequency-domain representation extraction network CFRNet, the image information can be converted to the frequency domain, the high-frequency components related to the rain streak noise in the image can be identified, and the low-frequency components related to the overall structure and background of the image can be retained, effectively extracting the robust latent frequency-domain representation.
[0065] S3. Perform dual-frequency latent representation enhancement on the latent frequency-domain representation to obtain the enhanced representation.
[0066] In the embodiments of the present invention, the dual-frequency latent representation enhancement is divided into a low-frequency latent representation enhancement branch and a high-frequency latent representation enhancement branch. By enhancing the representations of different frequencies separately, it effectively focuses on the low-frequency local texture details and object contours in the complex dynamic images of drones in rainy days, filters out the high-frequency rain noise, and thus provides a robust latent frequency-domain representation for the subsequent image representation rain removal and restoration process.
[0067] Specifically, the target of the low-frequency latent representation enhancement module is the information-rich low-frequency components. It efficiently enhances the low-frequency information by using the linear complexity and powerful information modeling ability of the frequency-domain state space module, thereby balancing performance and computational cost, which is beneficial for edge deployment. The high-frequency latent representation enhancement module uses the multi-head attention mechanism to focus on the distribution area of high-frequency rain noise, ensuring that the subsequent image representation rain removal and restoration process can accurately locate the area to be de-rained.
[0068] In the embodiments of the present invention, dual-frequency latent representation enhancement is performed on the latent frequency-domain representation to obtain an enhanced representation, including: performing low-frequency enhancement on the latent frequency-domain representation using the low-frequency latent representation enhancement branch in the pre-constructed dual-frequency latent representation enhancement network to obtain a low-frequency enhanced representation; performing high-frequency enhancement on the latent frequency-domain representation using the high-frequency latent representation enhancement branch in the dual-frequency latent representation enhancement network to obtain a high-frequency enhanced representation; and performing convolutional connection on the low-frequency enhanced representation and the high-frequency enhanced representation to obtain the enhanced representation.
[0069] In detail, the dual-frequency latent representation enhancement network can be expressed as:
[0070]
[0071] Among them, represents the latent frequency-domain representation, represents the layer normalization operation, represents the frequency-domain state space module, represents the residual wavelet transform convolution, represents the 1×1 convolution, represents the linear mapping, represents the multi-head self-attention, represents the connection operation in the channel dimension, represents the enhanced representation. It should be noted that the parameters not explained in the above formula are all intermediate parameters used in the calculation process.
[0072] In the embodiments of the present invention, the Dual-Frequency Latent-Representation Enhancement Network (DFRNet) can enhance the latent frequency-domain representation from both low-frequency and high-frequency aspects, thereby effectively focusing on the low-frequency local texture details and object contours in the complex dynamic images of drones in rainy days, and improving the effect of subsequent image restoration.
[0073] S4. Perform representation aggregation image restoration according to the enhanced representation to obtain a rain-removed image of the complex dynamic video image of the drone in rainy days.
[0074] In the embodiments of the present invention, the aggregation image restoration is to restore the image to the original size, so as to achieve a robust rain-removing effect for the complex dynamic images of drones in rainy days. The present invention can utilize the deep frequency-domain representation aggregation image restoration network, which mainly consists of a parallel spatial latent frequency-domain representation branch and a channel latent frequency-domain representation branch. Specifically, the spatial latent frequency-domain representation branch processes the frequency-domain spatial information through operations such as the above-mentioned frequency-domain state space module, max pooling, and average pooling, effectively restoring the background information disturbed by rain noise. At the same time, the channel latent frequency-domain representation branch processes the frequency-domain channel information through operations such as the above-mentioned frequency-domain state space module, global max pooling, and global average pooling, enhancing the expression of key original information in the image channel dimension, suppressing the expression of rain noise channel dimension information, and effectively fusing the results of the two. Finally, the image is restored to the original size through a multi-layer residual wavelet transform convolution, so as to achieve a robust rain-removing effect for the complex dynamic images of drones in rainy days.
[0075] Specifically, performing representation aggregation image restoration according to the enhanced representation to obtain a rain-removed image of the complex dynamic video image of the drone in rainy days includes: performing spatial representation aggregation on the enhanced representation using the spatial latent frequency-domain representation branch in the pre-constructed deep frequency-domain representation aggregation image restoration network to obtain a spatial aggregation representation; performing frequency-domain representation aggregation on the enhanced representation using the channel latent frequency-domain representation branch in the deep frequency-domain representation aggregation image restoration network to obtain a frequency-domain aggregation representation; performing a multi-layer residual wavelet transform convolution on the spatial aggregation representation and the frequency-domain aggregation representation to obtain the rain-removed image corresponding to the complex dynamic video image of the drone in rainy days.
[0076] Specifically, the deep frequency-domain representation aggregation image restoration network can be expressed as:
[0077]
[0078] Among them, represents the enhanced representation, represents the residual wavelet transform convolution, represents the frequency-domain state space module, represents the max pooling operation, represents an average pooling operation, represents a concatenation operation in the channel dimension, represents a 1×1 convolution, represents a global maximum pooling operation, represents a global average pooling operation, represents a multi-layer residual wavelet transform convolution, represents a de-rained image. It should be noted that the parameters not defined in the above formula are all intermediate parameters used in the calculation process.
[0079] Specifically, the deep frequency-domain feature aggregation image restoration network can be as Figure 8 shown in Figure 8 where represents a residual wavelet transform convolution, represents a frequency-domain state space module, represents a maximum pooling operation, represents an average pooling operation, represents a concatenation operation in the channel dimension, represents a 1×1 convolution, represents a global maximum pooling operation, represents a global average pooling operation, represents a multi-layer residual wavelet transform convolution, represents an element-wise multiplication operation, represents a tensor addition operation. In the figure, represents a channel dimension concatenation and convolution dimension reduction operation. The number of layers of the multi-layer residual wavelet transform convolution is adaptively determined according to the number of units in the contrast perception-based potential frequency-domain feature extraction network.
[0080] Specifically, the present invention combines an image contrast regularization prior module, a contrast perception-based potential frequency-domain feature extraction network, a dual-frequency potential feature enhancement network, and a deep frequency-domain feature aggregation image restoration network through a UAV rainy-day complex dynamic video image de-raining framework based on potential frequency-domain features, as specifically shown in Figure 9 as follows.
[0081] Among them, the image contrast regularization prior module (Image Contrast -Regularization PriorModule, ICPM) can fully utilize the maximum and minimum values of the channels of UAV rainy-day complex dynamic video images to calculate the image contrast, and extract the prominent regions of the image detail texture through regularization operations and convolution operations, so as to provide image contrast prior information for potential frequency-domain feature extraction, ensure that the image maintains a reasonable contrast and visual structure during the processing, avoid over-smoothing or information loss, effectively retain the key original representations of the image, and achieve a clearer and more natural image restoration effect;
[0082] Contrast-Aware Frequency-Domain Latent-Representation Extraction Network (CFRNet), which can fully extract the latent frequency-domain representation of complex dynamic video images of drones in rainy days, supplement the image contrast information at the same time, avoid problems such as excessive background smoothing or distortion caused by the high mobility of drones during the process of separating the high-frequency detailed texture information and low-frequency background information of the image, suppress the generation of artifacts, so that the object contours in the image are clearer, maintain the naturalness and layering of the image and significantly improve the visual effect;
[0083] Dual-Frequency Latent-Representation Enhancement Network (DFRNet), which can fully focus on the high-frequency detailed texture and low-frequency overall structure of complex dynamic video images of drones in rainy days through high-frequency and low-frequency dual-frequency representation modeling, and perform adaptive fusion and enhancement in the frequency domain, so as to effectively filter out high-frequency rain noise and retain the key original image information, and then accurately locate the rain-removing area of complex dynamic video images of drones in rainy days for challenges such as the high mobility of drones and the generally low resolution of objects in the image;
[0084] Deep Frequency-Domain Representation Aggregation Image-Restoration Network (DRANet), which can make full use of parallel processing of spatial frequency domain and channel frequency domain to focus on the spatial structure and texture information of complex dynamic video images of drones in rainy days, strengthen or suppress the representation contributions of different channels in the frequency domain, so as to effectively remove high-frequency rain noise while enhancing the local representation and overall structure of the image, optimize the network representation ability, and finally achieve a robust rain-removing effect for complex dynamic video images of drones in rainy days.
[0085] Therefore, through the present invention, the image clarity of drones in complex dynamic environments in rainy days can be effectively improved, problems such as raindrop occlusion, image blur, background blurring, and color distortion can be solved, the rain noise can be effectively removed for challenges such as the high dynamicity of drones and the generally low resolution of objects in the image by using the image frequency-domain information, and more effective detailed texture information and object contour information can be formed in the image representation, so as to achieve a robust and efficient video image rain-removing effect in complex dynamic scenes of drones in rainy days, finally improve the visual effect of the image, and obtain a more accurate rain-removed image.
[0086] Preferably, the method of the present invention can be used in an industrial drone edge vision camera deployed on an embedded GPU platform to achieve edge computing. Specifically, the edge computing implemented by the method of the present invention refers to collecting complex dynamic image data in rainy scenarios through an on-board edge vision camera on the drone flight body, and independently implementing edge vision intelligence in the camera, that is, obtaining high-precision, strong robustness, and real-time image de-raining results based on the above-mentioned method for removing rain from drone video images based on potential frequency domain representation at the image data acquisition end. The method of the present invention can serve the obtained image de-raining results in aspects such as fast human-drone interaction, communication data optimization, real-time response operation, intelligent analysis application, privacy protection, and data security. Ensure that the number of communications and communication volume with the cloud platform are reduced as much as possible, thereby reducing waiting time and computing costs. Given that the task of removing rain from complex dynamic images of drones in rainy days is deployed to the image data acquisition end, it can effectively reduce the congestion of the backbone network, alleviate bandwidth occupancy, achieve low latency, improve processing efficiency, speed up response requests, and further improve the quality of removing rain from complex dynamic images of drones in rainy days. In particular, the speed after edge deployment of the method of the present invention can reach the real-time processing speed, and its de-raining effect and image quality meet the requirements of industrial applications.
[0087] As Figure 10 shown, an embodiment of the present invention provides a functional module diagram of a device for removing rain from drone video images based on potential frequency domain representation.
[0088] The device 100 for removing rain from drone video images based on potential frequency domain representation of the present invention can be installed in an electronic device. According to the functions implemented, a device 100 for removing rain from drone video images based on potential frequency domain representation can include an image contrast regularization processing module 101, a potential frequency domain representation extraction module 102, a dual-frequency potential representation enhancement module 103, and an image restoration module 104. The modules of the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.
[0089] In this embodiment, the functions of each module / unit are as follows: The image contrast regularization processing module 101 is used to obtain complex dynamic video images of drones in rainy days, perform image contrast regularization processing on the complex dynamic video images of drones in rainy days, and obtain a contrast-regularized image; the potential frequency domain representation extraction module 102 is used to extract the potential frequency domain representation of the complex dynamic video images of drones in rainy days according to the contrast-regularized image to guide a pre-constructed potential frequency domain representation extraction network; the dual-frequency potential representation enhancement module 103 is used to perform dual-frequency potential representation enhancement on the potential frequency domain representation to obtain an enhanced representation; the image restoration module 104 is used to perform representation aggregation image restoration according to the enhanced representation to obtain a de-rained image of the complex dynamic video images of drones in rainy days.
[0090] Specifically, in one embodiment, each module in a drone video image de-raining device 100 based on latent frequency domain representation uses the same technical means as a drone video image de-raining method based on latent frequency domain representation in the accompanying drawings when in use, and can produce the same technical effects, which will not be elaborated here.
[0091] For example, although not shown, the electronic device may further include a power source (such as a battery) for powering each component. Preferably, the power source can be logically connected to at least one processor through a power management device, so as to implement functions such as charging management, discharging management, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, and a power status indicator. The electronic device may also include a variety of sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.
[0092] It should be understood that the embodiments are only for illustrative purposes and are not limited by this structure in the scope of the patent application.
[0093] The present invention also provides a computer-readable storage medium storing a computer program, which when executed by a processor, can implement a drone video image de-raining method according to any of the above embodiments based on latent frequency domain representation.
[0094] It should be noted that the computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium may include: any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM, Read-Only Memory). In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules is only a logical function division, and there may be other division methods in actual implementation.
[0095] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0096] In addition, in each embodiment of the present invention, each functional module can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.
[0097] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.
[0098] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
[0099] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0100] In addition, obviously, the term "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices recited in the system claims can also be implemented by one unit or device through software or hardware. The terms "first", "second", etc. are used to denote names and do not denote any particular order.
[0101] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for removing rain from UAV video images based on latent frequency domain representation, characterized in that, The method includes: Obtaining a complex dynamic video image of a drone in rainy weather, performing image contrast regularization processing on the complex dynamic video image of the drone in rainy weather to obtain a contrast-regularized image; Guided by the contrast-regularized image, extracting a latent frequency domain representation of the complex dynamic video image of the drone using a pre-constructed latent frequency domain representation extraction network; Performing dual-frequency latent representation enhancement on the latent frequency domain representation to obtain an enhanced representation; Performing representation aggregation image restoration based on the enhanced representation to obtain a rain-removed image of the complex dynamic video image of the drone in rainy weather; The performing image contrast regularization processing on the complex dynamic video image of the drone in rainy weather to obtain a contrast-regularized image includes: Using a pre-constructed image contrast regularization prior module to extract the contrast of the complex dynamic video image of the drone in rainy weather; Performing contrast regularization operation on the complex dynamic video image of the drone in rainy weather according to the contrast to obtain a regularized contrast; Performing residual wavelet transform convolution on the regularized contrast to obtain a contrast-regularized image; The image contrast regularization prior module can be expressed as: Among them, represents the complex dynamic video image of the drone in rainy days, represents function, represents the minimum value of the image channel of the complex dynamic video image of the drone in rainy days, represents the maximum value of the image channel of the complex dynamic video image of the drone in rainy days, represents the contrast regularization operation, represents the 1×1 convolution, represents the residual wavelet transform convolution, represents the contrast regularization image, 、 、 are all intermediate parameters used in the calculation process.
2. The method for removing rain from UAV video images based on potential frequency domain representation according to claim 1, wherein The residual wavelet transform convolution is expressed as: Among them, represents the input tensor of the residual wavelet transform convolution, represents the wavelet transform convolution, is a 1×1 convolution, represents the phantom convolution, represents the output tensor of the residual wavelet transform convolution, is an intermediate parameter used in the calculation process.
3. The method for removing rain from UAV video images based on potential frequency domain representation according to claim 1, characterized in that, The guiding the pre-constructed latent frequency domain representation extraction network by the contrast-regularized image to extract the latent frequency domain representation of the complex dynamic video image of the drone in rainy weather includes: Using the frequency domain state space module in the latent frequency domain representation extraction network to extract the image frequency domain representation in the complex dynamic video image of the drone in rainy weather; Using the frequency domain representation mining module in the latent frequency domain representation extraction network to calculate the spectral robust latent representation of the complex dynamic video image of the drone in rainy weather; Using the contrast prior refinement module in the latent frequency domain representation extraction network to perform representation mining on the contrast-regularized image to obtain a mined representation; Using the prior gating feed-forward module in the latent frequency domain representation extraction network to perform gating learning on the spectral robust latent representation and the mined representation to obtain a gated representation; Generating the latent frequency domain representation of the complex dynamic video image of the drone in rainy weather according to the image frequency domain representation and the gated representation.
4. The method for removing rain from UAV video images based on potential frequency domain representation according to claim 3, wherein, The frequency domain state space module can be expressed as: Among them, represents the complex dynamic video image of the drone in rainy days, represents the layer normalization operation, represents the linear mapping, represents the two-dimensional fast Fourier transform, represents the residual wavelet transform convolution, represents the GeLU activation function, represents the inverse two-dimensional fast Fourier transform, represents the 2D selective scanning module, represents the image frequency domain representation, , , are all intermediate parameters used in the calculation process.
5. The method for removing rain from UAV video images based on potential frequency domain representation according to claim 3, characterized in that, The frequency domain representation mining module can be expressed as: Among them, represents the complex dynamic video image of the drone in rainy days, represents the residual wavelet transform convolution, represents the separation operation, represents the asymmetric convolution, represents the ReLU activation function, represents the concatenation operation in the channel dimension, represents the 1×1 convolution, represents the two-dimensional fast Fourier transform, represents the GeLU activation function, represents the inverse two-dimensional fast Fourier transform, represents the spectrum-robust latent representation, , , , , are all intermediate parameters used in the calculation process.
6. The method for removing rain from UAV video images based on potential frequency domain representation according to claim 3, wherein The contrast prior refinement module can be expressed as: Among them, represents the contrast-regularized image, represents the residual wavelet transform convolution, represents the asymmetric convolution, represents function, represents the 1×1 convolution, represents the soft pooling operation, represents the mining feature, 、 、 are all intermediate parameters used in the calculation process.
7. The method for removing rain from UAV video images based on potential frequency domain representation according to claim 3, wherein The prior gating feed-forward module can be expressed by the formula: Among them, represents the spectrum-robust latent representation, represents the mined representation, represents the residual wavelet transform convolution, represents the asymmetric convolution, represents the Tanh activation function, represents the 1×1 convolution, represents the gated unit, represents the gated representation, 、 、 、 are all intermediate parameters used in the calculation process.
8. The method for removing rain from UAV video images based on potential frequency domain representation according to claim 1, wherein, The performing dual-frequency latent representation enhancement on the latent frequency domain representation to obtain an enhanced representation includes: Using the low-frequency latent representation enhancement branch in the pre-constructed dual-frequency latent representation enhancement network to perform low-frequency enhancement on the latent frequency domain representation to obtain a low-frequency enhanced representation; Using the high-frequency latent representation enhancement branch in the dual-frequency latent representation enhancement network to perform high-frequency enhancement on the latent frequency domain representation to obtain a high-frequency enhanced representation; Performing convolutional connection on the low-frequency enhanced representation and the high-frequency enhanced representation to obtain an enhanced representation.
9. The method for removing rain from UAV video images based on potential frequency domain representation according to claim 8, wherein The dual-frequency latent representation enhancement network can be expressed as: Among them, represents the potential frequency-domain representation, represents the layer normalization operation, represents the frequency-domain state space module, represents the residual wavelet transform convolution, represents the 1×1 convolution, represents the linear mapping, represents the multi-head self-attention, represents the concatenation operation on the channel dimension, represents the enhanced representation, , , are all intermediate parameters used in the calculation process.
10. A method for removing rain from UAV video images based on potential frequency domain representation according to claim 1, characterized in that, The performing representation aggregation image restoration based on the enhanced representation to obtain a rain-removed image of the complex dynamic video image of the drone in rainy weather includes: Aggregate the spatial latent frequency domain representation in the spatial latent frequency domain representation branch of the pre - constructed deep frequency domain representation aggregation image restoration network for the enhanced representation to obtain a spatial aggregation representation; Aggregate the frequency domain representation of the enhanced representation using the channel latent frequency domain representation branch in the deep frequency domain representation aggregation image restoration network to obtain a frequency domain aggregation representation; Perform multi - layer residual wavelet transform convolution on the spatial aggregation representation and the frequency domain aggregation representation to obtain the rain - removed image corresponding to the complex dynamic video image of the drone in rainy days.
11. A method for removing rain from UAV video images based on potential frequency domain representation according to claim 10, characterized in that, The deep frequency domain representation aggregation image restoration network can be expressed as: Among them, represents enhanced representation, represents residual wavelet transform convolution, represents a frequency-domain state space module, represents a max pooling operation, represents an average pooling operation, represents a concatenation operation in the channel dimension, represents a 1×1 convolution, represents a global max pooling operation, represents a global average pooling operation, represents multi-layer residual wavelet transform convolution, represents a de-rained image, , , , , , , are all intermediate parameters used in the calculation process.
12. An unmanned aerial vehicle video image de-raining device based on latent frequency domain representation, characterized in that, The device includes: An image contrast regularization processing module, which is used to obtain the complex dynamic video image of the drone in rainy days, extract the contrast of the complex dynamic video image of the drone in rainy days using the pre - constructed image contrast regularization prior module; perform contrast regularization operation on the complex dynamic video image of the drone in rainy days according to the contrast to obtain a regularized contrast; perform residual wavelet transform convolution on the regularized contrast to obtain a contrast - regularized image; A latent frequency domain representation extraction module, which is used to guide the pre - constructed latent frequency domain representation extraction network to extract the latent frequency domain representation of the complex dynamic video image of the drone in rainy days according to the contrast - regularized image; A dual - frequency latent representation enhancement module, which is used to perform dual - frequency latent representation enhancement on the latent frequency domain representation to obtain an enhanced representation; An image restoration module, which is used to perform representation aggregation image restoration according to the enhanced representation to obtain the rain - removed image of the complex dynamic video image of the drone in rainy days; The image contrast regularization prior module can be expressed as: Among them, represents the complex dynamic video image of the drone in rainy weather, represents a function, represents the minimum value of the image channel of the complex dynamic video image of the drone in rainy weather, represents the maximum value of the image channel of the complex dynamic video image of the drone in rainy weather, represents the contrast regularization operation, represents a 1×1 convolution, represents the residual wavelet transform convolution, represents the contrast regularization image, , , are all intermediate parameters used in the calculation process.