Multi-modal image defogging method and device, electronic equipment and storage medium

By extracting and fusing features from RGB images and radar point cloud data, and introducing the feature enhancement module of Vision Transformer, the problem of insufficient image dehazing accuracy in existing technologies is solved, and high-precision image restoration is achieved in complex haze environments.

CN121032837AActive Publication Date: 2025-11-28CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511563765.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2025-11-28
Estimated Expiration
2045-10-30

AI Technical Summary

Technical Problem

Existing image dehazing methods cannot meet the actual needs in the face of complex and changing hazy environments. Traditional methods make overly idealistic assumptions, while deep learning methods are limited by training data bias.

Method used

A multimodal image dehazing method is adopted, which extracts features from RGB images and radar point cloud data respectively, performs feature fusion through a parallel convolution module, and introduces a feature enhancement module based on Vision Transformer to construct an image dehazing model.

Benefits of technology

It improves image dehazing accuracy, better adapts to scenes with different haze densities and non-uniform haze distribution, and enhances image processing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121032837A_ABST
    Figure CN121032837A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal image defogging method and device, electronic equipment and a storage medium, and relates to the technical field of image processing, and the method comprises the steps: carrying out the feature extraction of an RGB image and radar point cloud data, and obtaining RGB image features and radar point cloud features; inputting into a first parallel convolution module and a second parallel convolution module for convolution processing to respectively obtain a first feature fusion result and a second feature fusion result; performing feature fusion through the first feature fusion result and the second feature fusion result to obtain a multi-modal image fusion result; based on a visual processing improved algorithm, a multi-modal image fusion result is input into an image defogging model to obtain a fused defogged image, the image defogging model comprises a target encoder, a target decoder and a feature enhancement module, and the feature enhancement module is constructed based on Vision Transform. According to the invention, the image defogging precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to a multi-modal image defogging method and device, electronic equipment and storage medium. BACKGROUND

[0002] Image defogging is of great significance in many fields, and can solve the image degradation problem under various complex atmospheric conditions such as non-uniform fog, thick fog, and haze-fog mixture, and restore a clear image. Traditional methods based on physical models usually need to estimate the assumptions of the fog and haze scene, which are often too idealized and cannot cope with the complex and changing fog and haze environment in practice. Although deep learning methods partially alleviate these problems, they are still limited by training data bias, and the image defogging precision cannot meet the actual needs. SUMMARY

[0003] The problem solved by the present application is how to improve the image defogging precision.

[0004] To solve the above problems, the present application provides a multi-modal image defogging method, device, electronic equipment and storage medium.

[0005] In a first aspect, the present application provides a multi-modal image defogging method, comprising: performing feature extraction on the RGB image and the radar point cloud data respectively to obtain RGB image features and radar point cloud features; inputting the RGB image features and the radar point cloud features into a first parallel convolution module and a second parallel convolution module for convolution processing to obtain first feature fusion results and second feature fusion results respectively, wherein the first parallel convolution module is used for matrix subtraction operation, and the second parallel convolution module is used for matrix addition operation; performing feature fusion through the first feature fusion results and the second feature fusion results to obtain multi-modal image fusion results; inputting the multi-modal image fusion results into an image defogging model based on a visual processing improved algorithm to obtain a fusion defogging image, wherein the image defogging model comprises a target encoder, a target decoder and a feature enhancement module, and the feature enhancement module is constructed based on a Vision Transformer.

[0006] Optionally, the construction process of the image defogging model comprises: inputting historical multi-modal image fusion results into an initial encoder for feature enhancement to obtain an initial encoder image; training the feature enhancement module through the initial encoder image to obtain a visual processing image; inputting the visual processing image into an initial decoder for image restoration to obtain an initial decoder image; perform mean square error loss function calculation according to the initial encoder image and the initial decoder image to obtain a total loss value; adjust model parameters of the initial encoder and the initial decoder according to the total loss value until the total loss value meets a preset condition, and take the initial encoder and the initial decoder after parameter adjustment as the target encoder and the target decoder, respectively.

[0007] Optionally, the initial encoder image includes an input image, a first convolution layer processed image and a second convolution layer processed image, the initial decoder image includes a third convolution layer processed image, a fourth convolution layer processed image and an output image, and the performing mean square error loss function calculation according to the initial encoder image and the initial decoder image to obtain a total loss value includes: perform mean square error loss function calculation on the input image and the output image to obtain a main loss value; perform mean square error loss function calculation on the first convolution layer processed image and the fourth convolution layer processed image to obtain a regularization term loss value; perform mean square error loss function calculation on the second convolution layer processed image and the third convolution layer processed image to obtain a constraint term loss value; obtain the total loss value according to the main loss value, the regularization term loss value and the constraint term loss value; wherein the total loss value includes: , wherein, the total loss value is, the main loss value is, the regularization term loss value is, the constraint term loss value is.

[0008] Optionally, the training the feature enhancement module through the initial encoder image to obtain a visual processing image includes: obtain a decomposition image by decomposing the initial encoder image; obtain the visual processing image according to the decomposition image based on a multi-head attention mechanism.

[0009] Optionally, the feature extraction is performed on the RGB image and the radar point cloud data respectively to obtain an RGB image feature and a radar point cloud feature, and the feature extraction includes: obtain the RGB image and the radar point cloud data; input the RGB image into an image feature extraction module to obtain the RGB image feature, wherein the image feature extraction module includes an image convolution layer and an image ReLU activation function layer. inputting the radar point cloud data into a radar point cloud feature extraction module to obtain the radar point cloud feature, wherein the radar point cloud feature extraction module comprises a radar point cloud convolution layer, a radar point cloud normalization layer and a radar point cloud ReLU activation function layer.

[0010] Optionally, the RGB image feature and the radar point cloud feature are inputted into a first parallel convolution module and a second parallel convolution module for convolution processing to obtain a first feature fusion result and a second feature fusion result, respectively, comprising: The RGB image feature and the radar point cloud feature are inputted into the first parallel convolution module for convolution processing to obtain the first feature fusion result, wherein the first parallel convolution module comprises a first RGB convolution branch and a first radar point cloud convolution branch. The first feature fusion result comprises: , wherein, the first feature fusion result is the output result of the RGB image feature after being processed by the first RGB convolution branch is the output result of the radar point cloud feature after being processed by the first radar point cloud convolution branch is The RGB image feature and the radar point cloud feature are inputted into the second parallel convolution module for convolution processing to obtain the second feature fusion result, wherein the second parallel convolution module comprises a second RGB convolution branch and a second radar point cloud convolution branch. The second feature fusion result is: , wherein, the second feature fusion result is the output result of the RGB image feature after being processed by the second RGB convolution branch is the output result of the radar point cloud feature after being processed by the second radar point cloud convolution branch is

[0011] Optionally, the first feature fusion result and the second feature fusion result are fused to obtain a multi-modal image fusion result, comprising: The first feature fusion result and the second feature fusion result are fused after being regularized to obtain the multi-modal image fusion result. The multi-modal image fusion result is: , Wherein, the Fusion_feature is the multi-modal image fusion result, is the first feature fusion result, is the second feature fusion result, Conv() is a convolution processing, BN() is a normalization processing, and ReLU() is a ReLU activation function processing.

[0012] In a second aspect, the present application provides a multi-modal image defogging device, comprising: a feature extraction module, configured to perform feature extraction on an RGB image and radar point cloud data respectively to obtain an RGB image feature and a radar point cloud feature; a convolution processing module, configured to input the RGB image feature and the radar point cloud feature into a first parallel convolution module and a second parallel convolution module for convolution processing to obtain a first feature fusion result and a second feature fusion result respectively, wherein the first parallel convolution module is configured to perform a matrix subtraction operation, and the second parallel convolution module is configured to perform a matrix addition operation; an image fusion module, configured to perform feature fusion through the first feature fusion result and the second feature fusion result to obtain a multi-modal image fusion result; an image defogging module, configured to input the multi-modal image fusion result into an image defogging model based on a visual processing improvement algorithm to obtain a fusion defogging image, wherein the image defogging model comprises a target encoder, a target decoder and a feature enhancement module, and the feature enhancement module is constructed based on a Vision Transformer.

[0013] In a third aspect, the present application provides an electronic device, comprising a memory and a processor; the memory, configured to store a computer program; the processor, configured to implement the multi-modal image defogging method of the first aspect when executing the computer program.

[0014] In a fourth aspect, the present application provides a computer readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the multi-modal image defogging method of the first aspect is implemented.

[0015] The multi-modal image defogging method, device, electronic equipment and storage medium have the beneficial effects that feature extraction is respectively performed on the RGB image and the radar point cloud data, the contour information of the image can be better acquired by extracting the features of the two modalities respectively, the difference features are captured by performing a matrix subtraction operation after the feature convolution processing of the two modalities by the first parallel convolution module, and the common features are strengthened by performing a matrix addition operation after the feature convolution processing of the two modalities by the second parallel convolution module. The features processed by the first parallel convolution module and the second parallel convolution module are fused to obtain a multi-modal image fusion result, and the features after fusion are enhanced, thereby laying a foundation for subsequent image processing. The multi-modal image fusion result is input into an image defogging model based on the improved algorithm of visual processing to obtain a fusion defogging image, and the image defogging is realized by a target encoder, a target decoder and a feature enhancement module. The feature enhancement module constructed based on the Vision Transformer is introduced, so that the image defogging model can better adapt to scenes with different fog density and non-uniform fog distribution, and the image defogging precision is improved. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1 A flowchart of a multi-modal image defogging method according to an embodiment of the present application; Figure 2 A structure diagram of an image defogging model according to an embodiment of the present application; Figure 3 A structure diagram of an image feature extraction module according to an embodiment of the present application; Figure 4 A structure diagram of a radar point cloud feature extraction module according to an embodiment of the present application; Figure 5 A structure diagram of a multi-modal image defogging device according to an embodiment of the present application; Figure 6 A structure diagram of an electronic equipment according to an embodiment of the present application. DETAILED DESCRIPTION

[0017] In order to make the above objectives, characteristics and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the drawings. Although some embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms, and should not be interpreted as being limited to the embodiments described herein, on the contrary, these embodiments are provided to make the present application more thorough and complete. It should be understood that the drawings and embodiments of the present application are only for illustrative purposes, and are not intended to limit the protection scope of the present application.

[0018] It should be understood that each of the steps recited in the method embodiments of the present application can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit performing the steps shown. The scope of the present application is not limited in this respect.

[0019] The term "comprises" and variations thereof used in the present document such as "comprising" and "comprises" are open-ended, that is, "comprising but not limited to"; the term "based on" is "based, at least in part, on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments"; the term "optionally" means "optional embodiment". Related definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc. mentioned in the present application are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0020] It should be noted that the modification of "one" or "multiple" mentioned in the present application is illustrative and not limiting, and those skilled in the art should understand that unless otherwise explicitly indicated in the context, it should be understood as "one or more".

[0021] The names of the messages or information exchanged between the devices in the embodiments of the present application are only for illustrative purposes, and are not used to limit the scope of the messages or information.

[0022] As shown in Figure 1 The multi-modal image defogging method provided by the embodiment of the present application comprises: In step 110, feature extraction is performed on the RGB image and the radar point cloud data respectively to obtain the RGB image feature and the radar point cloud feature.

[0023] Specifically, for different sizes of input images, first, size adaptation processing is performed, which adjusts the image to 256x256 pixels by center cropping or random cropping, so that the final input size is 256x256x3 RGB image. Subsequently, the image data is normalized, and the original pixel value in the range of 0-255 is divided by 255, so that all pixel values are normalized to the interval [0, 1]. This processing method can not only meet the model input requirements, but also improve the numerical stability of the training process. For the matching radar point cloud information, the X, Y coordinate data (corresponding to the image pixel position) and the Intensity (intensity) and Distance (distance) two key parameters are directly extracted from the Excel table. These data are organized into a two-channel three-dimensional matrix, where the first channel stores the intensity value and the second channel stores the distance value, and the matrix size is XxY. We normalize the data in each channel respectively, and the specific method is to divide the channel value by the maximum value in the channel, so that all values are mapped to the range [0, 1]. Finally, to maintain consistency with the image data, the three-dimensional matrix is also randomly cropped to adjust its size to 256x256x2, ensuring that the sizes of the multi-modal data are matched.

[0024] Step 120, input the RGB image features and the radar point cloud features into a first parallel convolution module and a second parallel convolution module for convolution processing, respectively obtaining a first feature fusion result and a second feature fusion result, wherein the first parallel convolution module is used for matrix subtraction operation, and the second parallel convolution module is used for matrix addition operation.

[0025] Specifically, because the RGB image and the radar point cloud information are different data information, multi-modal feature fusion is needed, so in the implementation process of the feature fusion module, a two-channel network structure is adopted to process and cross-fuse the RGB image features and the radar point cloud features during fusion, which is divided into a first parallel convolution module Diff (difference branch) and a second parallel convolution module Sum (sum branch), which respectively capture difference features and strengthen common features.

[0026] Step 130, performing feature fusion through the first feature fusion result and the second feature fusion result to obtain a multi-modal image fusion result.

[0027] Specifically, the results obtained from the two branches are respectively convolved and regularized, and then the two matrices of the same size obtained are added, so that the final multi-modal image fusion result obtained contains the pixel features of the original foggy image of the RGB image and the intensity and distance information of the radar point cloud data, that is, the final result contains pixel features and contour features.

[0028] In step 140, the multi-modal image fusion result is input into an image defogging model based on a visual processing improvement algorithm to obtain a fusion defogging image, wherein the image defogging model comprises a target encoder, a target decoder and a feature enhancement module, and the feature enhancement module is constructed based on a Vision Transformer.

[0029] Specifically, in the foregoing steps, the extraction and enhancement of feature information have been realized, the multi-modal image fusion result is input into an image defogging model to obtain a fusion defogging image, and a single image defogging method based on detail enhancement convolution and content guided attention is adopted. The design of the image defogging model is that a feature enhancement module constructed based on a Vision Transformer is inserted between a target encoder and a target decoder. In the feature enhancement module, a self-attention mechanism is used to realize efficient local information aggregation. The Vision Transformer (ViT) is a model structure that applies the classic Transformer architecture in natural language processing to computer vision tasks.

[0030] In this embodiment, feature extraction is performed on the RGB image and the radar point cloud data respectively. By extracting the features of the two modalities respectively, the contour information of the image can be better obtained. After the features of the two modalities are convoluted by the first parallel convolution module and the second parallel convolution module, matrix subtraction and matrix addition operations are performed respectively, so as to capture the difference features and strengthen the common features. The features processed by the first parallel convolution module and the second parallel convolution module are fused to obtain a multi-modal image fusion result, and the features after fusion are enhanced to lay a foundation for subsequent image processing. The multi-modal image fusion result is input into an image defogging model based on a visual processing improvement algorithm to obtain a fusion defogging image, and the image defogging is realized by a target encoder, a target decoder and a feature enhancement module. The introduction of the feature enhancement module constructed based on the Vision Transformer enables the image defogging model to better adapt to scenes with different fog densities and non-uniform fog distributions, thereby improving the image defogging precision.

[0031] Optionally, the construction process of the image defogging model comprises: inputting a historical multi-modal image fusion result into an initial encoder for feature enhancement to obtain an initial encoder image; training the feature enhancement module through the initial encoder image to obtain a visual processing image; inputting the visual processing image into an initial decoder for image restoration to obtain an initial decoder image; calculating a mean square error loss function according to the initial encoder image and the initial decoder image to obtain a total loss value; adjusting model parameters of the initial encoder and the initial decoder according to the total loss value until the total loss value meets a preset condition, and taking the initial encoder and the initial decoder after parameter adjustment as the target encoder and the target decoder respectively.

[0032] Specifically, as shown in Figure 2 The initial encoder (Encoder) is responsible for extracting features from the input image and performing multi-scale feature representation. The initial decoder (Decoder) gradually restores the features extracted by the initial encoder to the original image size and generates the final output. The output of the initial encoder is sent to the feature enhancement module (Vision Transformer), which is processed by the feature enhancement module and then sent to the initial decoder to obtain the final result.

[0033] In this optional embodiment, image defogging is realized through the feature enhancement module to improve the accuracy of image defogging. The mean square error is used for loss function calculation, which measures the average square difference between the predicted value and the true value, facilitating the use of gradient descent and other optimization algorithms, especially suitable for image reconstruction tasks such as image defogging.

[0034] Optionally, the initial encoder image includes an input image, a first convolutional layer processed image, and a second convolutional layer processed image, the initial decoder image includes a third convolutional layer processed image, a fourth convolutional layer processed image, and an output image, and the mean square error loss function calculation according to the initial encoder image and the initial decoder image obtains a total loss value, including: performing mean square error loss function calculation on the input image and the output image to obtain a main loss value; performing mean square error loss function calculation on the first convolutional layer processed image and the fourth convolutional layer processed image to obtain a regularization term loss value; performing mean square error loss function calculation on the second convolutional layer processed image and the third convolutional layer processed image to obtain a constraint term loss value; obtaining the total loss value according to the main loss value, the regularization term loss value, and the constraint term loss value; wherein the total loss value includes: , wherein, the total loss value, the main loss value, the regularization term loss value, the constraint term loss value.

[0035] In some more specific embodiments, as Figure 2As shown, the initial encoder (Encoder) part first inputs the historical multi-modal image fusion result into the input layer (fusion_feature input) of the initial encoder to obtain an input image, and obtains a first convolution layer processing image after the input image is processed by a first convolution layer (DEB N1). The first convolution layer processing image is processed by a down-sampling layer (Down-sampling), and the processed image is input into a second convolution layer (DEB N2) to obtain a second convolution layer processing image. The second convolution layer processing image is processed by a down-sampling layer and input into a feature enhancement module.

[0036] The initial decoder (Decoder) part processes the image processed by the feature enhancement module by an up-sampling layer (Up-sampling), inputs the processed image into a third convolution layer to obtain a third convolution layer (DEB N3) processing image, fuses the third convolution layer processing image and the image processed by the down-sampling layer in the initial encoder (Encoder) part with respect to the first convolution layer processing image by a Fusion module, processes the fused image by an up-sampling layer, inputs the processed image into a fourth convolution layer (DEB N4) to obtain a fourth convolution layer processing image. The fourth convolution layer processing image is input into the output layer (output) of the initial decoder to obtain an output image. The convolution layer includes a convolution block (Conv) and a feature enhancement module.

[0037] Specifically, three loss functions are designed for calculating loss values in different processing processes. The three loss functions all use mean square error (MSE) to calculate the loss function, and the loss function calculation process includes: , wherein, is, is the value of the matrix obtained by processing the corresponding decoder, is the value obtained by processing the corresponding encoder, and N is the number of current training samples.

[0038] The three loss values are respectively assigned different weights. The weights of the main loss value Loss_1, the regularization term loss value Loss_2 and the constraint term loss value Loss_3 are 0.6, 0.2 and 0.2 respectively. Therefore, the corresponding total loss value includes: , wherein, is the total loss value, is the main loss value, is the regularization term loss value, is the constraint term loss value.

[0039] According to the above loss, the parameters of the neural network are adjusted by gradient descent using the SGD optimizer, and the parameters of the network are continuously adjusted according to the loss, so that the loss is gradually reduced, and finally when the loss is stable at a very low value, it is considered that the model training has converged. SGD (Stochastic Gradient Descent) is one of the most basic and widely used optimization algorithms in deep learning. It iteratively updates model parameters to minimize the loss function. For the setting of the optimization parameter, the learning rate is set to 0.001, and during the training process, the learning rate can be adjusted. There needs to be a suitable learning rate, too high, which can easily lead to model divergence, the value is too high, the model converges too slowly.

[0040] , where, is the updated model parameter for the next iteration, is the learning rate, is the model parameter at the current iteration time, is the model parameter.

[0041] In this optional embodiment, three loss functions are designed to calculate the error of different processing stages of the model. The three loss functions act on different levels of the encoder. This multi-level supervision mechanism can resist feature distortion and improve overall robustness. Multiple supervision signals avoid gradient disappearance, making it easier to train deep networks.

[0042] Optionally, the training of the feature enhancement module through the initial encoder image to obtain a visual processing image comprises: obtaining a decomposition image by decomposing the initial encoder image; based on the multi-head attention mechanism, obtaining the visual processing image according to the decomposition image.

[0043] In some more specific embodiments, based on the visual processing improvement on the traditional transformer model, the Vision Transformer is used to process the fused data processed by the encoder. The input initial encoder image, i.e. the feature matrix, is divided into several non-overlapping small matrices. The input feature matrix is 64*64*3, which is decomposed into 64 8*8*3 small matrices (patches). Each small matrix (patch) is flattened into a vector, which includes: , where, is the vector corresponding to the i-th small matrix, Linear() is a linear transformation layer, Flatten() is a flattening operation, N is the number of image blocks, and the value of dimension D is the number of matrix elements (i.e. 8*8*3=192 dimensions). Because the processing process is out-of-order processing, position information is added to each vector to form the final token. The position information records the position of the current vector corresponding to the small matrix in the original large matrix, which is used for subsequent small matrix splicing. The process includes: , wherein, is the token corresponding to the i-th patch, is the position information corresponding to the i-th patch. The final token sequence is After obtaining the token sequence, three independent weight matrices are randomly generated. According to the token sequence, K, Q, and V are generated. The attention value of each token is calculated according to the obtained K, Q, and V. The process includes: , , wherein, is the transpose matrix of matrix K, is a scaling factor, and Softmax() is a Softmax function.

[0044] According to the token output from the encoder, the patch block corresponding to the current token is processed, and the model only pays attention to the current input patch and is not affected by other patches. Then the content of the next patch block is predicted to speed up the processing of each patch. This process is implemented through Masked Multi-Head Self-Attention. The process includes: , wherein, M and V are related matrices. After processing each patch, each token is converted from a one-dimensional vector to a three-dimensional tensor. Finally, according to the position information added in the encoding process, it is spliced into a three-dimensional matrix as it was originally transmitted into the module for further processing. Cross Entropy is used to calculate the loss, including: , wherein, is the loss, is the input feature matrix, is the feature matrix output after the module processing, c is the number of different categories, after each batch of training is implemented, according to the loss value, all parameters of the model are updated through back propagation According to the chain rule, gradient descent is performed to update the parameters. Including: , , , By continuously updating these parameters, the feature processing capability of the model is continuously optimized.

[0045] Optionally, the feature extraction is performed on the RGB image and the radar point cloud data respectively to obtain the RGB image feature and the radar point cloud feature, including: Obtaining the RGB image and the radar point cloud data; Inputting the RGB image into an image feature extraction module to obtain the RGB image feature, wherein the image feature extraction module includes an image convolution layer and an image ReLU activation function layer; Inputting the radar point cloud data into a radar point cloud feature extraction module to obtain the radar point cloud feature, wherein the radar point cloud feature extraction module includes a radar point cloud convolution layer, a radar point cloud normalization layer and a radar point cloud ReLU activation function layer.

[0046] In some specific embodiments, as shown in Figure 3 The image feature extraction module is implemented using convolution operation, including four convolution layers of different sizes, and an image ReLU activation function layer is inserted after the last convolution layer to increase the nonlinearity of the model. The ReLU activation function includes: , The parameters used by each convolution layer are kernel=3 (convolution kernel size is 3*3), stride=1 (convolution kernel moves by 1 each time), and padding='same' (automatic edge padding). By setting the same convolution parameters, the size of the image feature information matrix after extraction is 128*128*3. The specific convolution operation output size calculation formula is as follows: , , Among them, is the height of the output after convolution operation, is the height of the input matrix of the convolution operation, is the width of the output matrix of the convolution operation, ​P is the width of the input matrix for the convolution operation, P is the number of padding pixels, K is the kernel size, and S is the kernel stride.

[0047] In some specific embodiments, such as Figure 4 As shown, similar to the image feature extraction module described above, the radar point cloud feature extraction module also primarily extracts feature information through convolution operations. It includes three convolutional layers of different sizes, and after each convolutional operation, a RescaleNorm layer is performed for layer normalization. An image ReLU activation function layer is interspersed after the last convolutional layer to prevent excessively large feature values ​​from affecting subsequent convolutional operations. The formula used for the RescaleNorm operation is as follows: , Where x is the input tensor, For the set scalar, These are the parameters for the entire module model. These are the parameters of the upper convolutional layer. For setting parameter weights, This is the scaling value. This is the offset. Since the size of the input point cloud matrix information matrix is ​​the same as that of the foggy image matrix (except that it has one less channel), the parameters used in the convolution operation in this module are the same as those used in the convolution operation for feature extraction from the foggy image. The final output radar point cloud feature matrix size is 128*128*2.

[0048] In this optional embodiment, RGB images provide rich texture and color information, making them suitable for recognizing the appearance features of objects. Radar point cloud data provides high-precision spatial geometric information (such as distance and shape), is unaffected by illumination, and is suitable for detecting the position and contours of objects. ReLU is proposed as an activation function, making it easier to inversely recover images during image dehazing.

[0049] Optionally, the step of inputting the RGB image features and the radar point cloud features into a first parallel convolution module and a second parallel convolution module for convolution processing to obtain a first feature fusion result and a second feature fusion result, respectively, includes: The RGB image features and the radar point cloud features are input into the first parallel convolution module for convolution processing to obtain the first feature fusion result. The first parallel convolution module includes a first RGB convolution branch and a first radar point cloud convolution branch. The first feature fusion result includes: , in, is the first feature fusion result, is an output result of the RGB image feature after being processed by the first RGB convolution branch, is an output result of the radar point cloud feature after being processed by the first radar point cloud convolution branch; The RGB image feature and the radar point cloud feature are input into the second parallel convolution module for convolution processing to obtain a second feature fusion result, wherein the second parallel convolution module includes a second RGB convolution branch and a second radar point cloud convolution branch. The second feature fusion result is: , wherein, is the second feature fusion result, is an output result of the RGB image feature after being processed by the second RGB convolution branch, is an output result of the radar point cloud feature after being processed by the second radar point cloud convolution branch.

[0050] Optionally, the feature fusion through the first feature fusion result and the second feature fusion result to obtain a multi-modal image fusion result includes: regularization processing and then feature fusion through the first feature fusion result and the second feature fusion result to obtain the multi-modal image fusion result; The multi-modal image fusion result is: , wherein, Fusion_feature is the multi-modal image fusion result, is the first feature fusion result, is the second feature fusion result, Conv() is convolution processing, BN() is normalization processing, and ReLU() is ReLU activation function processing.

[0051] In some more specific embodiments, the first parallel convolution module Diff (difference branch) and the second parallel convolution module Sum (sum branch) are used to process the fog image feature information and the radar point cloud feature information, respectively. In the Diff branch, the preprocessed fog image feature information and the radar point cloud feature information are matrices of 128*128*3 and 128*128*2, respectively. The fog image feature information data is input into the first small branch of the Diff branch as input, the image feature information is enhanced, and the linear connection is realized through the convolution and ReLU module, then the result after each convolution operation is recorded, and after the last convolution operation, the results obtained by each convolution operation are added to obtain the final enhanced image feature information, and the obtained feature information data is a matrix of 64*64*3; the radar point cloud feature information is input into the second small branch of the Diff branch, and the same processing as the fog image feature information is performed, and finally the enhanced radar point cloud feature is obtained, and the obtained features are matrices of 64*64*2. Then, the obtained feature values are subtracted, and the specific operation is that the value of each channel matrix of the image feature information is subtracted from the sum of the values of the two channel matrices of the radar point cloud feature matrix. Finally, the processing result of the Diff branch is obtained. The first parallel convolution module includes a first RGB convolution branch and a first radar point cloud convolution branch, the first RGB convolution branch includes: , , , wherein, is the kth convolution layer, is the output result of the i-th convolution layer of the first RGB convolution branch, i is 1, 2, 3. The first radar point cloud convolution branch includes: , , , wherein, is the output result of the i-th convolution layer of the first radar point cloud convolution branch, and the final difference calculation includes: .

[0052] In the Sum branch, the operation is similar to that of the Diff branch, except that the last subtraction absolute value operation is changed to an addition operation. The two results obtained by the convolution operation of the two small branches are respectively denoted as and , and the final result obtained by the branch includes: The results of the two branches are respectively convolved and regularized, and then two matrices of the same size obtained are added, so that the final matrix contains the pixel features of the original foggy image and the intensity and distance information of the radar point cloud, that is, the final result contains the pixel features and contour features of the original image.

[0053] As shown in Figure 5 The embodiment of the application provides a multi-modal image defogging device, which comprises: A feature extraction module 10 is configured to extract features from an RGB image and radar point cloud data respectively to obtain RGB image features and radar point cloud features. A convolution processing module 20 is configured to input the RGB image features and the radar point cloud features into a first parallel convolution module and a second parallel convolution module for convolution processing to obtain first feature fusion results and second feature fusion results respectively, wherein the first parallel convolution module is configured to perform a matrix subtraction operation, and the second parallel convolution module is configured to perform a matrix addition operation. An image fusion module 30 is configured to perform feature fusion through the first feature fusion results and the second feature fusion results to obtain multi-modal image fusion results. An image defogging module 40 is configured to input the multi-modal image fusion results into an image defogging model based on a visual processing improvement algorithm to obtain a fusion defogging image, wherein the image defogging model comprises a target encoder, a target decoder and a feature enhancement module, and the feature enhancement module is constructed based on a Vision Transformer.

[0054] The multi-modal image defogging device of the embodiment is used to implement the multi-modal image defogging method as described above, and has the same advantages as the multi-modal image defogging method compared with the prior art, which will not be described here again.

[0055] Optionally, the multi-modal image defogging device further comprises a model construction module, which is configured to input historical multi-modal image fusion results into an initial encoder for feature enhancement to obtain an initial encoder image. The feature enhancement module is trained through the initial encoder image to obtain a visual processing image. The visual processing image is input into an initial decoder for image restoration to obtain an initial decoder image. A mean square error loss function is calculated according to the initial encoder image and the initial decoder image to obtain a total loss value. Adjusting model parameters of the initial encoder and the initial decoder according to the total loss value until the total loss value meets a preset condition, and taking the initial encoder and the initial decoder after parameter adjustment as the target encoder and the target decoder respectively.

[0056] Optionally, the model construction module is further configured to: perform mean square error loss function calculation on the input image and the output image to obtain a main loss value. Optionally, the model construction module is further configured to: perform mean square error loss function calculation on the image processed by the first convolutional layer and the image processed by the fourth convolutional layer to obtain a regularization term loss value. Optionally, the model construction module is further configured to: perform mean square error loss function calculation on the image processed by the second convolutional layer and the image processed by the third convolutional layer to obtain a constraint term loss value. Optionally, the model construction module is further configured to: obtain the total loss value according to the main loss value, the regularization term loss value and the constraint term loss value. Optionally, the total loss value includes: , Optionally, the total loss value includes:

[0057] Optionally, the model construction module is further configured to: obtain a decomposed image by decomposing the initial encoder image. Optionally, the model construction module is further configured to: obtain the visual processing image according to the decomposed image based on a multi-head attention mechanism.

[0058] Optionally, the feature extraction module 10 is specifically configured to: obtain the RGB image and the radar point cloud data. Optionally, the feature extraction module 10 is specifically configured to: input the RGB image into an image feature extraction module to obtain the RGB image feature, wherein the image feature extraction module includes an image convolutional layer and an image ReLU activation function layer. Optionally, the feature extraction module 10 is specifically configured to: input the radar point cloud data into a radar point cloud feature extraction module to obtain the radar point cloud feature, wherein the radar point cloud feature extraction module includes a radar point cloud convolutional layer, a radar point cloud normalization layer and a radar point cloud ReLU activation function layer.

[0059] Optionally, the convolutional processing module 20 is specifically configured to: input the RGB image feature and the radar point cloud feature into the first parallel convolutional module for convolutional processing to obtain the first feature fusion result, wherein the first parallel convolutional module includes a first RGB convolutional branch and a first radar point cloud convolutional branch. Optionally, the first feature fusion result includes: ​​​​ , wherein, is the first feature fusion result, is an output result of the RGB image feature after being processed by the first RGB convolution branch, is an output result of the radar point cloud feature after being processed by the first radar point cloud convolution branch; input the RGB image feature and the radar point cloud feature into the second parallel convolution module for convolution processing to obtain the second feature fusion result, wherein the second parallel convolution module comprises a second RGB convolution branch and a second radar point cloud convolution branch; wherein, the second feature fusion result is: , wherein, is the second feature fusion result, is an output result of the RGB image feature after being processed by the second RGB convolution branch, is an output result of the radar point cloud feature after being processed by the second radar point cloud convolution branch.

[0060] Optionally, the image fusion module 30 is specifically configured to perform feature fusion after regularization processing of the first feature fusion result and the second feature fusion result to obtain the multi-modal image fusion result. wherein, the multi-modal image fusion result is: , wherein, Fusion_feature is the multi-modal image fusion result, is the first feature fusion result, is the second feature fusion result, Conv() is convolution processing, BN() is normalization processing, and ReLU() is ReLU activation function processing.

[0061] As shown in Figure 6 , the embodiment of the present application provides an electronic device 600, which comprises a memory 610 and a processor 620; the memory 610 is used for storing a computer program; the processor 620 is used for implementing the multi-modal image defogging method as described above when the computer program is executed.

[0062] Alternatively, an electronic device 600 comprises a memory 610 and a processor 620 coupled to the memory 610; the memory 610 is configured to store a computer program; the processor 620 is configured to perform the following operations when the computer program is executed: Feature extraction is performed on the RGB image and the radar point cloud data respectively to obtain RGB image features and radar point cloud features; The RGB image features and the radar point cloud features are input into a first parallel convolution module and a second parallel convolution module for convolution processing to obtain a first feature fusion result and a second feature fusion result respectively, wherein the first parallel convolution module is used for performing a matrix subtraction operation, and the second parallel convolution module is used for performing a matrix addition operation. Feature fusion is performed through the first feature fusion result and the second feature fusion result to obtain a multi-modal image fusion result. The multi-modal image fusion result is input into an image defogging model based on a visual processing improvement algorithm to obtain a fusion defogging image, wherein the image defogging model comprises a target encoder, a target decoder and a feature enhancement module, and the feature enhancement module is constructed based on a Vision Transformer.

[0063] The embodiment of the present application provides a computer readable storage medium, and the storage medium stores a computer program.

[0064] Alternatively, a non-volatile computer readable storage medium stores a computer program. Feature extraction is performed on the RGB image and the radar point cloud data respectively to obtain RGB image features and radar point cloud features; The RGB image features and the radar point cloud features are input into a first parallel convolution module and a second parallel convolution module for convolution processing to obtain a first feature fusion result and a second feature fusion result respectively, wherein the first parallel convolution module is used for performing a matrix subtraction operation, and the second parallel convolution module is used for performing a matrix addition operation. Feature fusion is performed through the first feature fusion result and the second feature fusion result to obtain a multi-modal image fusion result. The multi-modal image fusion result is input into an image defogging model based on a visual processing improvement algorithm to obtain a fusion defogging image, wherein the image defogging model comprises a target encoder, a target decoder and a feature enhancement module, and the feature enhancement module is constructed based on a Vision Transformer.

[0065] An electronic device 600, which can be a server or a client of the present application, will now be described, which is an example of a hardware device that can be applied to aspects of the present application. The electronic device 600 is intended to represent various forms of digital electronic computer devices such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device 600 can also represent various forms of mobile devices such as personal digital assistants, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present application described and / or claimed in this document.

[0066] The electronic device 600 includes a computing unit that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) or a computer program loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The computing unit, the ROM, and the RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.

[0067] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a computer readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments. The storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM), or a random access memory (RAM). In this application, the units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, i.e., they can be located in one place or distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiments of the present application. In addition, the functional units in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically, or two or more units can be integrated into one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0068] Although the present application is disclosed as above, the protection scope of the present application is not limited to this. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present application, and these changes and modifications will fall within the protection scope of the present application.

Claims

1. A multimodal image dehazing method, characterized in that, include: Feature extraction was performed on RGB images and radar point cloud data respectively to obtain RGB image features and radar point cloud features; The RGB image features and the radar point cloud features are input into a first parallel convolution module and a second parallel convolution module for convolution processing to obtain a first feature fusion result and a second feature fusion result, respectively. The first parallel convolution module is used to perform matrix subtraction operations, and the second parallel convolution module is used to perform matrix addition operations. By fusing the first feature fusion result and the second feature fusion result, a multimodal image fusion result is obtained; Based on the improved visual processing algorithm, the multimodal image fusion result is input into the image dehazing model to obtain the fused dehazed image. The image dehazing model includes a target encoder, a target decoder, and a feature enhancement module, and the feature enhancement module is built based on Vision Transformer.

2. The multimodal image dehazing method according to claim 1, characterized in that, The process of constructing the image dehazing model includes: The fusion results of historical multimodal images are input into the initial encoder for feature enhancement to obtain the initial encoder image; The feature enhancement module is trained using the initial encoder image to obtain a visually processed image; The visually processed image is input into the initial decoder for image reconstruction to obtain the initial decoder image; The total loss value is obtained by calculating the mean squared error loss function based on the initial encoder image and the initial decoder image; The model parameters of the initial encoder and the initial decoder are adjusted according to the total loss value until the total loss value meets the preset conditions. The initial encoder and the initial decoder after parameter adjustment are then used as the target encoder and the target decoder, respectively.

3. The multimodal image dehazing method according to claim 2, characterized in that, The initial encoder image includes an input image, a first convolutional layer processed image, and a second convolutional layer processed image. The initial decoder image includes a third convolutional layer processed image, a fourth convolutional layer processed image, and an output image. The calculation of the mean squared error loss function based on the initial encoder image and the initial decoder image to obtain the total loss value includes: The main loss value is obtained by calculating the mean squared error loss function using the input image and the output image; The mean squared error loss function is calculated by processing the image through the first convolutional layer and the image through the fourth convolutional layer to obtain the regularization term loss value; The mean squared error loss function is calculated by processing the image through the second convolutional layer and the image through the third convolutional layer to obtain the constraint term loss value; The total loss value is obtained based on the main loss value, the regularization term loss value, and the constraint term loss value. The total loss value includes: , in, The total loss value, The main loss value, The loss value of the regularization term. The loss value is the constraint term.

4. The multimodal image dehazing method according to claim 2, characterized in that, The step of training the feature enhancement module using the initial encoder image to obtain a visually processed image includes: The decomposed image is obtained by decomposing the initial encoder image; The visually processed image is obtained from the decomposed image based on the multi-head attention mechanism.

5. The multimodal image dehazing method according to claim 1, characterized in that, The step of extracting features from RGB images and radar point cloud data respectively to obtain RGB image features and radar point cloud features includes: Acquire the RGB image and the radar point cloud data; The RGB image is input into the image feature extraction module to obtain the RGB image features, wherein the image feature extraction module includes an image convolutional layer and an image ReLU activation function layer; The radar point cloud data is input into the radar point cloud feature extraction module to obtain the radar point cloud features. The radar point cloud feature extraction module includes a radar point cloud convolutional layer, a radar point cloud normalization layer, and a radar point cloud ReLU activation function layer.

6. The multimodal image dehazing method according to claim 1, characterized in that, The step of inputting the RGB image features and the radar point cloud features into a first parallel convolution module and a second parallel convolution module for convolution processing, and obtaining a first feature fusion result and a second feature fusion result respectively, includes: The RGB image features and the radar point cloud features are input into the first parallel convolution module for convolution processing to obtain the first feature fusion result. The first parallel convolution module includes a first RGB convolution branch and a first radar point cloud convolution branch. The first feature fusion result includes: , in, The result of the first feature fusion. The output result of the RGB image features after processing by the first RGB convolution branch. The output result of the radar point cloud features after processing by the first radar point cloud convolutional branch; The RGB image features and the radar point cloud features are input into the second parallel convolution module for convolution processing to obtain the second feature fusion result. The second parallel convolution module includes a second RGB convolution branch and a second radar point cloud convolution branch. The result of the second feature fusion is as follows: , in, This is the result of the second feature fusion. The output result of the RGB image features after processing by the second RGB convolution branch. The output result is the radar point cloud features processed by the second radar point cloud convolution branch.

7. The multimodal image dehazing method according to claim 6, characterized in that, The step of fusing features using the first feature fusion result and the second feature fusion result to obtain a multimodal image fusion result includes: The multimodal image fusion result is obtained by performing regularization processing on the first feature fusion result and the second feature fusion result; The multimodal image fusion result is as follows: , Wherein, Fusion_feature is the result of the multimodal image fusion. The result of the first feature fusion. The result of the second feature fusion is represented by Conv(), which is a convolution process, BN() is a normalization process, and ReLU() is a ReLU activation function process.

8. A multimodal image dehazing device, characterized in that, include: The feature extraction module is used to extract features from RGB images and radar point cloud data respectively, to obtain RGB image features and radar point cloud features; The convolution processing module is used to input the RGB image features and the radar point cloud features into the first parallel convolution module and the second parallel convolution module for convolution processing, and obtain the first feature fusion result and the second feature fusion result respectively. The first parallel convolution module is used to perform matrix subtraction operation, and the second parallel convolution module is used to perform matrix addition operation. The image fusion module is used to perform feature fusion using the first feature fusion result and the second feature fusion result to obtain a multimodal image fusion result; The image dehazing module is used to input the multimodal image fusion result into the image dehazing model to obtain a fused dehazed image based on the improved visual processing algorithm. The image dehazing model includes a target encoder, a target decoder, and a feature enhancement module, and the feature enhancement module is built based on Vision Transformer.

9. An electronic device, characterized in that, Including memory and processor; The memory is used to store computer programs; The processor is configured to implement the multimodal image dehazing method as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the multimodal image dehazing method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-modal fusion method for SAR and visible light image feature enhancement

    CN120032216A

  • Multi-modal image fusion method, system, medium and equipment

    CN120107084A

  • Image dehazing method, apparatus and device

    WO2023040462A1