A multi-spectral prior-guided tone mapping method for mobile phone images

By installing a multi-spectral sensor on the mobile terminal, the spectrum-perceptual self-attention module is used to fuse the feature information of multi-spectral and RGB images to generate spectrum enhancement bilateral grid coefficients, solving the problem of insufficient accuracy of mobile phone image tone enhancement in complex environments, and achieving better local color consistency and high dynamic range.

CN119693285BActive Publication Date: 2025-06-27NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510215713.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-27
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

The prior art is difficult to achieve accurate cell phone image tone enhancement in complex environments, especially in terms of image enhancement using additional spectral information.

Method used

By installing a multi-spectral sensor on the mobile terminal, the multi-spectral image and RGB image are obtained, the spectrum sensing self-attention module is used to fuse the characteristic information, generate the spectral enhancement bilateral grid coefficient, and generate the final tone enhancement output diagram through the guide map.

Benefits of technology

It is realized that multi-spectral information is effectively embedded without changing the existing image enhancement network structure to generate images with better local color consistency and stronger high dynamic range.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119693285B_ABST
    Figure CN119693285B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-spectral prior-guided mobile phone image tone mapping method, which includes the following steps: Step 1, an additional multi-spectral sensor is mounted on the mobile device to obtain the multi-spectral image of the scene in the visible light band, and at the same time, the visible light RGB image captured by the main camera of the mobile device is obtained; Step 2, after the multi-spectral image and the RGB image are subjected to feature extraction, they are input into the spectral-aware self-attention module through feature embedding to obtain the spectral-enhanced bilateral grid coefficients; Step 3, the RGB image and the multi-spectral image are jointly used as the guidance map, and the color transformation coefficients of each pixel point are obtained by interpolating in the spectral-enhanced bilateral grid coefficients through the guidance map, so as to obtain the final tone-enhanced output map. The present invention makes full use of the additional low-spatial-resolution multi-spectral image information for the tone enhancement task in mobile phone images, and overcomes the limited spectral imaging performance caused by the physical space limitation of the mobile device from the algorithm level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of spectral applications and image enhancement, and particularly relates to a multi-spectral prior-guided tone mapping method for mobile phone images. Background Art

[0002] In recent years, the progress of miniaturized spectrometers has enabled them to be integrated into mobile devices, thus broadening the scope of possible applications, such as medical diagnosis, food safety assessment, face anti-counterfeiting recognition, and color analysis. However, currently, spectral sensors in mobile phone photography are mainly used for light source estimation in automatic white balance (AWB) operations. Spectral information has the potential to enhance the aesthetic and perceptual quality of images. In mobile phone images, the goal of tone enhancement is to adjust the brightness, contrast, and color characteristics of the image to enhance the details in the highlight and shadow areas while preserving the natural appearance. However, accurate tone enhancement faces challenges in various complex environments. How to utilize additional spectral information to improve the image enhancement step, especially the tone enhancement process, in the mobile image signal processing (ISP) pipeline is an area that requires in-depth research. Summary of the Invention

[0003] Object of the Invention: The technical problem to be solved by the present invention is to provide a multi-spectral prior-guided tone mapping method for mobile phone images in view of the deficiencies of the prior art, thereby further improving the performance of mobile phone images.

[0004] The method of the present invention includes the following steps:

[0005] Step 1, an additional multi-spectral sensor is mounted on the mobile device, which usually requires a wider spectral band than a conventional RGB camera, to obtain a multi-spectral image I of the scene in the visible light band msi , and at the same time, obtain a visible light RGB image I taken by the main camera of the mobile device rgb . This step makes full use of the complementary characteristics of the RGB camera and the multi-spectral sensor in terms of spatial and spectral resolutions, and through appropriate parameter selection, an optimal multi-sensor complementary configuration is achieved;

[0006] Step 2, after the multi-spectral image I msi and the RGB image I rgb are subjected to feature extraction, they are input into the spectral perception self-attention module through feature embedding. The spectral perception self-attention module utilizes the complementary characteristics between I msi and I rgb to fuse the multi-spectral image I msi and the RGB image I rgb to obtain the spectral enhancement bilateral grid coefficient Φ(x, y, z); where x and y respectively represent the length and width in the spatial dimension of the bilateral grid coefficient, and z represents the guiding value;

[0007] Step 3: Jointly use the RGB image I rgb and the multispectral image I msi as the guidance map g, generate a guidance vector through the guidance map g, interpolate the color transformation coefficients of each pixel point in the spectral enhancement bilateral grid coefficient Φ(x, y, z), so as to obtain the final tone enhancement output map; using Φ(x, y, z) after multispectral enhancement, the generated image has better local color consistency.

[0008] Step 1 includes:

[0009] Collect the visible light RGB image I through the mobile phone main camera rgb , the height and width of the visible light RGB image I rgb are H and W respectively, and obtain the multispectral image I through the multispectral sensor msi , the height and width of the multispectral image I msi are h and W respectively;

[0010] The mobile phone main camera has a relatively high spatial resolution of H×W, and the multispectral sensor has a relatively high spectral resolution (L spectral channels, usually L is about 10) and a relatively low spatial resolution of h×w; it is reasonable that the size of H×W is X1 (usually taking values from 8 to 16) times the size of h×w;

[0011] In the hardware layout of the sensor, it is necessary to place the mobile phone main camera and the multispectral sensor as close as possible to reduce the parallax between the cameras, and set the optical axis distance between the mobile phone main camera and the multispectral sensor to be less than the threshold (usually 10 mm);

[0012] Multispectral image RGB image where L represents the number of spectral channels of the multispectral image; represents the real number space.

[0013] Step 2 includes:

[0014] Step 2-1: Scale the multispectral image I msi to the same spatial size as the RGB image I rgb :

[0015] Step 2-2: Perform feature extraction on the multispectral image I msi and the RGB image I rgb respectively to obtain the feature maps and the feature map where C represents the channel dimension size of the feature map, and M and N respectively represent the height and width of the feature map;

[0016] Step 2-3, concatenate the feature map and along the channel dimension of the feature map to obtain the concatenated fused feature map

[0017] Step 2-4, perform feature embedding mapping on the fused feature map to generate query vector Q, key vector K, and value vector V;

[0018] Step 2-5, input query vector Q, key vector K, and value vector V into the spectral perception self-attention module for processing: In the spectral perception self-attention module, query vector Q is reshaped into vector key vector K is reshaped into vector value vector V is reshaped into vector

[0019] Vector and The dot product result generates the spectral perception map Reweighted by the spectral perception map A, the formula is:

[0020]

[0021] where softmax represents the normalized exponential function and σ is a learnable scaling parameter; The spectral perception self-attention module fuses the cross-spectrum fusion features in a residual learning manner with the original feature map

[0022] is a 3×3 convolution used to adaptively weight the feature map ;

[0023] is a 3×3 convolution used to adaptively weight the feature map ;

[0024] is a 3×3 convolution for adaptively weighting ;

[0025] respectively represent the multi-spectral feature map and RGB feature map output after being processed by the spectral perception self-attention module;

[0026] Step 2-6, replace the multi-spectral image I and the RGB image I in Step 2-2 with the multi-spectral feature map msi and the RGB feature map rgb , repeat Step 2-2 to Step 2-5, and finally in the output feature map Based on this, the bilateral grid coefficients Φ(x, y, z) after spectral information enhancement are predicted through 1×1 convolution. Φ(x, y, z) contains the color transformation coefficients at the spatial position (x, y).

[0027] In step 2-2, the depth convolution is a 3*3 convolution.

[0028] In step 2-3, the 1×1 convolution W1 and the 3×3 convolution W3 are applied to the fused feature map to generate the query vector Q, the key vector K, and the value vector V:

[0029]

[0030] where respectively represent the 1×1 convolution W1 and the 3×3 convolution W3 acting on the query vector Q;

[0031] respectively represent the 1×1 convolution W1 and the 3×3 convolution W3 acting on the key vector K;

[0032] respectively represent the 1×1 convolution W1 and the 3×3 convolution W3 acting on the value vector V.

[0033] Step 3 includes:

[0034] Step 3-1, convert the RGB image I rgb and the multispectral image I msi into a guidance map g;

[0035] Step 3-2, generate a guidance vector through the guidance map g, and interpolate the transformation coefficients of each pixel point in the bilateral grid coefficients Φ(x, y, z);

[0036] Step 3-3, based on the transformation coefficients obtained in step 3-2, perform a color transformation operation on each pixel point of the input RGB image I rgb to obtain the final output image I target .

[0037] In step 3-1, the following formula is used to convert the RGB image I rgb and the multispectral image I msi into the guidance map g. The formula is:

[0038]

[0039] where and respectively represent the 1×1 convolution acting on the RGB image I rgb and the 1×1 convolution acting on the multispectral image Imsi 1×1 convolution, M 3×3 is a learnable matrix of size 3×3, b′ and b are learnable biases; c represents the channel dimension size.

[0040] In step 3-2, the transformation coefficients [a 11 , a 12 , a 13 ,..., a 43 are obtained using the following formula:

[0041] [a 11 , a 12 , a 13 ,..., a 43 = Φ(m, n, g(m, n)),

[0042] where (m, n) represents the coordinate point of a certain pixel in I rgb , m and n respectively represent the coordinate values in the image spatial dimension, g(m, n) represents the corresponding value at this coordinate position in the guidance map, a 11 , a 12 , a 13 ,..., a 43 represent the coefficients of the color transformation matrix, where a 11 represents the value in the first row and first column of the color transformation matrix, a 12 represents the value in the first row and second column of the color transformation matrix, and so on.

[0043] In step 3-3, the output image I target is obtained using the following formula:

[0044]

[0045] where v r , v g , v b respectively represent the pixel values of the red (r), green (g), and blue (b) channels of the input image I rgb at the coordinate point (m, n), v′ r , v′ g , v′ b respectively represent the pixel values of the red (r), green (g), and blue (b) channels of the output image I target at the coordinate point (m, n); for the color transformation matrix, a 33 represents the value in the third row and third column of the above color transformation matrix, and so on; by parallelizing the processing of all pixel values of the input image I rgb , the final output image I target can be obtained.

[0046] The present invention also provides an electronic device, including a processor and a memory, where the memory stores program code, and when the program code is executed by the processor, the processor is caused to execute the steps of the above-mentioned method.

[0047] The present invention also provides a storage medium storing a computer program or instruction, and when the computer program or instruction runs on a computer, the steps of the above-mentioned method are executed.

[0048] Generally speaking, through the utilization of multi-spectral prior and the above-mentioned technological innovations, including synergistically acquiring complementary RGB images and multi-spectral images, a tone mapping method guided by multi-spectral prior, and predicting a tone-enhanced target map based on a joint guidance map and spectral-enhanced bilateral grid coefficients, the present invention can effectively integrate additional spectral information into the tone enhancement task of mobile phone imaging.

[0049] The present invention has the following beneficial effects: (1) The present invention can make full use of additional multi-spectral image information with low spatial resolution for the tone enhancement task in mobile phone imaging, and overcome the limited spectral imaging performance caused by physical space limitations at the algorithm level for mobile devices;

[0050] (2) By utilizing the complementarity between RGB images with high spatial resolution and low spectral resolution and spectral images with low spatial resolution and high spectral resolution, and effectively integrating multi-sensor information through a tone mapping method guided by multi-spectral prior, more accurate spectral-enhanced bilateral grid coefficients for color mapping can be generated;

[0051] (3) The present invention can embed multi-spectral information in a plug-and-play manner without changing the main structure of the existing image enhancement network. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The following further specifically describes the present invention in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.

[0053] Figure 1 is a flowchart of the method of the present invention.

[0054] Figure 2 is an overall flowchart of tone enhancement guided by multi-spectral prior.

[0055] Figure 3 is a local high dynamic range comparison diagram between a high dynamic range enhancement model (HDRNet) and the method proposed by the present invention.

[0056] Figure 4 is a color comparison diagram between a high dynamic range enhancement model (HDRNet) and the method proposed by the present invention.

[0057] Figure 5 It is a contrast graph of the semantic consistency between the high-dynamic range enhancement model (HDRNet) and the method proposed in the present invention. Detailed implementation manners

[0058] As Figure 1 shown, an embodiment of the present invention provides a multi-spectral prior-guided tone mapping method for mobile phone images, including the following steps:

[0059] Step 1, an additional multi-spectral sensor is mounted on a mobile device (such as a smart phone), an appropriate spectral band is selected, and a multi-spectral image I of the scene in the visible light band is obtained through the multi-spectral sensor msi , and at the same time, an RGB (i.e., red, green, and blue) image I captured by the main camera of the mobile device is obtained rgb . The purpose of this step is to make full use of the high spatial resolution characteristics of the RGB camera and the high spectral resolution characteristics of the multi-spectral sensor, so as to make full use of the complementary characteristics of the RGB camera and the multi-spectral sensor, thereby overcoming the imaging limitations in the narrow physical space of the mobile device;

[0060] Step 2, a multi-spectral prior-guided tone mapping method is proposed, including image feature extraction, feature embedding mapping, and spectral-aware self-attention module. The multi-spectral image I msi and the RGB image I rgb are first converted into deep feature representations through feature extraction, and then input into the spectral-aware self-attention module through feature embedding mapping. By using the complementary characteristics between I msi and I rgb , the feature information from the dual-path sensors is fully fused in the spectral-aware self-attention module to obtain the spectral-enhanced bilateral grid coefficients Φ(x, y, z); Image feature extraction is to convert the input visible light image and multi-spectral image into deep feature representations, feature embedding mapping is to fuse and map the visible light and multi-spectral deep features to generate query vector Q, key vector K, and value vector V, and the spectral-aware self-attention module is to use the generated query vector Q, key vector K, and value vector V to generate more accurate spectral-enhanced bilateral grid color mapping coefficients;

[0061] Step 3, jointly using the RGB image I rgb and the multi-spectral image I msi as the guidance map g, generating a guidance vector through the guidance map g, interpolating in the bilateral grid coefficients Φ(x, y, z) to obtain the transformation coefficient of each pixel point, so as to obtain the final tone-enhanced output map.

[0062] Step 1 specifically includes the following steps:

[0063] Step 1-1: Since the physical space of the mobile device is limited, it is necessary to fully utilize the complementary characteristics of the existing visible light sensor and the multispectral sensor. Therefore, in the selection of sensors, the visible light sensor needs to have a high spatial resolution, while the multispectral sensor needs to have a high spectral resolution. In the spectral dimension, considering the spectral perception range of the silicon-based sensor, the spectral bands of the multispectral sensor are usually selected between 400 and 1000 nm, which has a wider spectral perception band than the conventional RGB camera; in the spatial dimension, it is a reasonable range that the spatial resolution of the RGB camera is 8 to 16 times that of the multispectral sensor.

[0064] Step 1-2: In the hardware layout of the sensors, the visible light sensor and the multispectral sensor need to be as close as possible to reduce the parallax between the cameras. To facilitate downstream vision task processing and reduce the difficulty of image alignment, the optical axis distance between the sensors generally needs to be less than 10 mm.

[0065] Step 1-3: Mark the multispectral image obtained by the above multi-sensor shooting as and the RGB image where L represents the number of spectral channels of the multispectral image, usually about 10, h and w represent the length and width in the spatial dimension of the multispectral image respectively; H and W represent the length and width in the spatial dimension of the RGB image respectively; among them, I msi has a small spatial resolution and a large spectral resolution, while I rgb has a large spatial resolution and a small spectral resolution, and the two form a complementary relationship in the spatial dimension and the spectral dimension.

[0066] Step 2 specifically includes the following steps:

[0067] Step 2-1: As Figure 1 shown, in order to ensure that the feature maps in the feature extraction network have the same size, first scale the multispectral image I msi to the same size as I rgb : Then use the multispectral prior-guided tone mapping method proposed by the present invention to enhance the images of the mobile device for the scaled multispectral image and the RGB image;

[0068] Step 2-2: For the multispectral image I msi and the visible light image Irg b First, perform consistent feature extraction operations through depth convolution. The feature extraction module usually consists of a cascade of 3*3 convolutions. Denote the feature maps obtained from the multispectral image and the visible light image based on the feature extraction module as and and Have the same feature map size, where H, W, and C represent the length, width, and number of channels of the depth feature map respectively;

[0069] Step 2-3, for the feature map and First, perform splicing at the channel level. Splicing is the simplest and most direct strategy for fusing feature maps of different modalities. Through the splicing operation, the fused feature map The number of channels of the feature map is and twice as much;

[0070] Step 2-4, perform feature embedding mapping on the feature map to generate the query vector Q (Query), key vector K (Key), and value vector V (Value). Use the 1×1 convolution W1 and 3×3 convolution W3 to act on the feature map to generate the query vector Q, key vector K, and value vector V. Specifically, represent the 1×1 convolution W1 and 3×3 convolution W3 acting on Q, K, and V respectively. Through this operation, it is convenient to fuse information from different spectral bands in the subsequent spectral perception self-attention module:

[0071]

[0072] Step 2-5, the query vector Q and key vector K are resized to a size of Therefore, the resized query vector and the key vector The dot product result generates the spectral perception map The spectral perception map A represents the mutual relationship between different spectral channels. The resized value vector The importance between different channels can be reweighted by the spectral perception map A, thereby modeling the correlation between the spectral information from the RGB camera and the multispectral sensor;

[0073] Step 2-6, the spectral perception self-attention module is expressed as:

[0074]

[0075] where σ is a learnable scaling parameter that controls the intensity of the dot product result. and are 3×3 convolutions used to perform adaptive weighting on and respectively, is a 3×3 convolution that performs adaptive weighting on respectively, Represent the feature maps of multispectral and RGB after being processed by the spectral perception self-attention module; the spectral perception self-attention module fuses cross-spectral fusion features in a residual learning manner with the original feature maps The spectral perception self-attention module fuses cross-spectral fusion features in a residual learning manner with the original feature maps Through the above design, more accurate color mapping coefficients can be learned in a progressive manner while fully utilizing the information from the dual-path sensors;

[0076] Step 2-7, input the output of the spectral perception self-attention module into the next layer for processing, and respectively perform feature extraction, feature embedding mapping, and spectral perception self-attention module operations. Repeat steps 2-2 to 2-7 in the above process operation, and predict the bilateral grid coefficients Φ(x, y, z) after spectral information enhancement through a 1×1 convolution on the finally output feature map Here, x and y represent the length and width in the spatial dimension of the bilateral grid coefficients, and z represents the guiding value. The color transformation parameters learned for a specific image enhancement task are stored in the bilateral grid coefficients;

[0077] Step 3 specifically includes the following steps:

[0078] Step 3-1, convert the RGB image I rgb and the multispectral image I msi into a guiding map g, where and respectively represent 1×1 convolutions, M 3×3 is a learnable matrix of size 3×3, b' and b are learnable biases. Through the combination of the multispectral image I msi and the RGB image I rgb a more accurate guiding map g can be generated:

[0079]

[0080] Step 3-2, perform interpolation on the spectral enhancement bilateral grid coefficients Φ(x, y) through the guiding map g generated by combining I rgb and I msi Specifically, determine the position of the interpolation points in the grid through trilinear interpolation, and calculate the transformation coefficients of each pixel point through comprehensive weighted calculation of the values and corresponding weights of the surrounding grid vertices;

[0081] Step 3-3, as Figure 2 shown, perform a transformation operation on each pixel point of the input RGB image I rgb to obtain the final output image Itarget . Specifically, for a certain coordinate point (m, n) in I rgb , where m and n respectively represent the coordinate values in the image spatial dimension, the pixel value of this coordinate point is [v r , v g , v b . The corresponding value g(m, n) of the guidance map g at this coordinate position is taken, and the coefficients [a 11 , a 12 , a 13 ,..., a 43 of the color transformation matrix are obtained from the bilateral grid coefficients Φ(x, y, z) through the guidance vector [m, n, g(m, n)], that is:

[0082] [a 11 , a 12 , a 13 ,..., a 43 = Φ(m, n, g(m, n)),

[0083] After obtaining the transformation coefficients, the pixel value conversion is achieved through the following formula:

[0084]

[0085] where v r , v g , v b respectively represent the pixel values of the red (r), green (g), and blue (b) channels of the input image I rgb at the coordinate point (m, n), and v′ r , v′ g , v′ b respectively represent the pixel values of the red (r), green (g), and blue (b) channels of the output image I target at the coordinate point (m, n); for the color transformation matrix, a 33 represents the value in the third row and third column of the above color transformation matrix, and so on; by parallelizing the pixel values of all input images I rgb , the final output image I target is obtained.

[0086] Step 3-4, as shown in Figure 3 , Figure 4 and Figure 5 , the first column is the 16-bit input image, the second column is the result processed by the high dynamic range enhancement model (HDRNet), and the third column is the result processed by the method of the present invention. Qualitatively analyzing the experimental results, it can be seen from Figure 3 that the method of the present invention has a stronger local high dynamic range, and from Figure 4It can be seen that the method proposed by the present invention has more accurate colors, from Figure 5 It can be seen that the method proposed by the present invention has better semantic consistency; the introduction of additional spectral information can bring advantages such as stronger local high dynamic range, more accurate colors, and semantic consistency.

[0087] As shown in the quantitative analysis experimental results in Table 1, in the hue enhancement task, the method proposed by the present invention is superior to the previous transform-based image enhancement methods in terms of indicators such as PSNR, SSIM, and ΔE * For example, the under-exposed image enhancement model (UPE), the three-dimensional lookup table-based models (3D LUT, 4D LUT, CLUT, SepLUT), the deep photo enhancement model (DPE), and the conditional sequence modulation image enhancement model (CSRNet). Ours in Table 1 is the method of the present invention.

[0088] Table 1

[0089]

[0090] The present invention provides a multi-spectral prior-guided mobile phone image tone mapping method. There are many methods and ways to specifically implement this technical solution. The above is only the preferred implementation mode of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and retouches can be made, and these improvements and retouches should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by existing technologies.

Claims

1. A multi-spectral prior-guided mobile phone image tone mapping method, characterized in that: The following steps are involved: Step 1: Equip the mobile terminal with an additional multispectral sensor to obtain a multispectral image of the scene in the visible light band. msi , and at the same time obtain the visible light RGB image I taken by the main camera of the mobile terminal rgb ; Step 2: Multispectral Image I msi With RGB image I rgb After feature extraction, the feature is embedded and input into the spectral perception self-attention module. The spectral perception self-attention module uses I msi with I rgb The complementary characteristics between them, fusion of multispectral images I msi With RGB image I rgb , obtain the spectral enhancement bilateral grid coefficient Φ(x,y,z); where x,y represent the length and width of the bilateral grid coefficient in the spatial dimension, respectively, and z represents the guide value; Step 3 includes: Step 3-1: transform the RGB image I rgb With multispectral image I msi Transformed into a guide graph g; Step 3-2, generate a guide vector through the guide map g, and interpolate the transformation coefficient of each pixel in the bilateral grid coefficient Φ(x, y, z); Step 3-3, based on the transformation coefficients obtained in step 3-2, the input RGB image I rgb Perform color transformation on each pixel to get the final output image I target ; In step 3-1, the RGB image I is converted into rgb With multispectral image I msi Transformed into the guide graph g, the formula is: in and Respectively represent the effect on RGB image I rgb The 1×1 convolution and the multispectral image I msi 1×1 convolution, M 3×3 is a learnable matrix of size 3×3, b′ and b are learnable biases; c represents the channel dimension size.

2. The method according to claim 1, characterized in that Step 1 includes: collecting visible light RGB image I through the main camera of the mobile terminal rgb , visible light RGB image I rgb The height and width of are H and W respectively, and the multispectral image I is obtained by the multispectral sensor. msi , multispectral image I msi The height and width of are h and w respectively; The size of H×W is X1 times the size of h×w; Set the optical axis distance between the mobile main camera and the multispectral sensor to be less than the threshold; Multispectral imagery RGB images Where L represents the number of spectral channels of the multispectral image; Represents the real number space.

3. The method according to claim 2, characterized in that Step 2 includes: Step 2-1: multispectral image I msi Scale to the same RGB image I rgb Same space size: Step 2-2: Deep convolution of multispectral image I msi With RGB image I rgb Perform feature extraction and obtain feature maps With feature map Where C represents the channel dimension size of the feature map, M and N represent the height and width of the feature map respectively; Step 2-3, the feature map and Splice in the feature map channel dimension to obtain the spliced ​​fusion feature map Step 2-4, fusion feature map Perform feature embedding mapping to generate query vector Q, key vector K and value vector V; Step 2-5, the query vector Q, key vector K and value vector V are input into the spectral-aware self-attention module for processing: In the spectral-aware self-attention module, the query vector Q is reshaped into a vector The key vector K is reshaped into a vector The value vector V is reshaped into a vector vector and The dot product of Reweighted by the spectral perception map A, the formula is: Where softmax represents the normalized exponential function, σ is a learnable scaling parameter; the spectral-aware self-attention module fuses cross-spectral fusion features in a residual learning manner Compared with the original feature map It is used to map the feature Perform adaptive weighted 3×3 convolution; It is used to map the feature Perform adaptive weighted 3×3 convolution; Yes Perform adaptive weighted 3×3 convolution; They represent the multi-spectral feature map and RGB feature map output after being processed by the spectral-aware self-attention module; Step 2-6: Multispectral feature map RGB feature map Replace the multispectral image I in step 2-2 msi With RGB image I rgb , after repeating steps 2-2 to 2-5, the feature map output at the end Based on this, a 1×1 convolution prediction is performed to obtain the bilateral grid coefficient Φ(x, y, z) enhanced by spectral information. Φ(x, y, z) contains the color transformation coefficient at the spatial position (x, y).

4. The method according to claim 3, characterized in that In step 2-2, the depth convolution is a 3*3 convolution.

5. The method according to claim 4, characterized in that In step 2-3, 1×1 convolution W1 and 3×3 convolution W3 are used to fusion feature map Generate query vector Q, key vector K and value vector V: in They represent the 1×1 convolution W1 and 3×3 convolution W3 acting on the query vector Q respectively; They represent the 1×1 convolution W1 and 3×3 convolution W3 acting on the key vector K respectively; They represent the 1×1 convolution W1 and 3×3 convolution W3 acting on the value vector V respectively.

6. The method according to claim 5, characterized in that In step 3-2, the coefficients of the color transformation matrix are obtained using the following formula [a 11 ,a 12 ,a 13 ,…,a 43 ]: [a 11 ,a 12 ,a 13 ,…,a 43 ]=Φ(m,n,g(m,n)), Where (m,n) represents I rgb The coordinate point of a pixel in the image, m, n represent the horizontal and vertical coordinates in the image space dimension respectively, g(m, n) represents the corresponding guidance value of the coordinate point (m, n) in the guidance map, a 11 ,a 12 ,a 13 ,…,a 43 Represents the coefficients of the color transformation matrix, where a 11 Represents the value of the first row and first column of the color transformation matrix, a 12 Represents the value of the first row and second column of the color transformation matrix, and so on; In step 3-3, the output image I is obtained using the following formula: target : where v r ,v g ,v b Represents the input image I rgb The pixel value of the red channel, the pixel value of the green channel, and the pixel value of the blue channel at the coordinate point (m, n), v r ′,v g ′,v′ b Represents the output graph I target The pixel value of the red channel, the pixel value of the green channel, and the pixel value of the blue channel at the coordinate point (m, n); For the color transformation matrix, a 33 Represents the value of the third row and third column in the above color transformation matrix, and so on; By processing all input images I in parallel rgb The pixel value of the final output image I is obtained target .

7. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 6.

8. A storage medium, characterized in that: A computer program or instruction is stored, and when the computer program or instruction is run on a computer, the steps of the method according to any one of claims 1 to 6 are executed.

Citation Information

Patent Citations

  • A multispectral image de-blurring method based on gradient domain priori

    CN109360161A

  • Multi-hyperspectral image fusion method guided by low-rank prior and spatial spectrum information

    CN114862731A