Image processing method and device

By using hyperspectral images to assist in denoising input images and fusing features with neural networks and attention mechanisms, the problem of poor image denoising performance in existing technologies is solved. This achieves enhancement of image texture and details, reduction of noise, and improvement of the signal-to-noise ratio of the image.

CN120876276APending Publication Date: 2025-10-31HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410548810.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-30
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing image denoising solutions have limited effectiveness in denoising images, especially in low-light or high-light scenes, which can easily lead to the loss of image details and textures, and cannot effectively improve the signal-to-noise ratio of the image.

Method used

The method utilizes the rich information of hyperspectral images to assist in denoising input images. It extracts features from hyperspectral images through neural networks and fuses them with the features of the input image. It also uses attention mechanisms and aberration estimation techniques to improve image quality.

Benefits of technology

It improves the signal-to-noise ratio of the image, enhances the texture and detail of the image, reduces noise, and improves the overall quality of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876276A_ABST
    Figure CN120876276A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method and device, and is used for assisting an input image in texture and detail enhancement by using a hyperspectral image, thereby achieving a denoising effect for the input image, and obtaining an output image with a high signal-to-noise ratio. Comprising the steps that firstly, an input image and a hyperspectral image are acquired, the input image is an image acquired by an image sensor, and spectral channels of the hyperspectral image are not completely identical to those of the input image, for example, the number of the channels of the hyperspectral image is larger than that of the channels of the input image; or the hyperspectral image comprises channels which are not included in the input image and the like, so that the hyperspectral image can comprise richer information or more accurate information compared with the input image; extracting features from the input image to obtain a first feature, and obtaining a second feature according to the features extracted from the hyperspectral image; and fusing the first feature and the second feature to obtain a fused feature, and then performing recovery according to the fused feature to obtain an output image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision, and more particularly to an image processing method and apparatus. Background Technology

[0002] Image sensors are sensors used in electronic devices to acquire images. The imaging process of image sensors is often affected by noise, which can originate from sources such as the imaging module or circuitry. Noise typically manifests in images as specular noise, blur, or artifacts, leading to a degraded image quality. For example, in some scenarios, such as low-light or bright-light scenes, details or textures may be lost. Image denoising methods aim to remove noise from images while preserving their original content. Common image denoising techniques utilize information from other sensors besides the image sensor to assist in image denoising. However, existing denoising aids have limited effectiveness; therefore, improving image denoising performance remains a crucial problem to be solved. Summary of the Invention

[0003] This application provides an image processing method and apparatus for enhancing the texture and details of an input image using hyperspectral images, thereby achieving a denoising effect on the input image and obtaining an output image with a high signal-to-noise ratio.

[0004] In view of the above, in a first aspect, this application provides an image processing method, comprising: first, acquiring an input image and a hyperspectral image, wherein the input image is an image acquired by an image sensor, and the spectral channels of the hyperspectral image are not exactly the same as those of the input image, for example, the number of channels in the hyperspectral image is greater than the number of channels in the input image, or the hyperspectral image includes channels not included in the input image, etc., which can be understood as the hyperspectral image including richer or more accurate information than the input image; then extracting features from the input image to obtain a first feature, and obtaining a second feature based on the features extracted from the hyperspectral image; fusing the first feature and the second feature to obtain a fused feature, and then recovering the output image based on the fused feature.

[0005] In this embodiment, the richer or more accurate information contained in the hyperspectral image can be used to assist in the enhancement or denoising of the input image, thereby enhancing the texture or details of the input image, reducing noise in the input image, and obtaining an output image with a high signal-to-noise ratio.

[0006] In one possible implementation, the aforementioned fusion of the first feature and the second feature to obtain the fused feature may include: fusing the first feature and the second feature based on an attention mechanism to obtain the fused feature.

[0007] In this embodiment, features extracted from the input image and the hyperspectral image can be fused based on an attention mechanism. This allows the attention mechanism to focus on the similarity between features, thereby enhancing the features extracted from the input image and improving the input image by reducing noise and obtaining an output image with a higher signal-to-noise ratio.

[0008] In one possible implementation, the aforementioned fusion of the first and second features to obtain the fused feature may further include: acquiring historical frame images, where the historical frame image is the previous frame image of the input image (e.g., if the input image is a frame in video data, the historical frame image can be the previous frame image of the input image in the video data); extracting features from the historical frame to obtain a third feature; and fusing the first, second, and third features to obtain the fused feature. In this embodiment, the historical frame image can be combined to further enhance the current frame, thereby further reducing noise in the current frame image.

[0009] In one possible implementation, obtaining the second feature based on features extracted from the hyperspectral image may include: filtering spectral channel information from the hyperspectral image to obtain a spectral channel reconstructed image; and obtaining the second feature based on features extracted from the spectral channel reconstructed image. In this embodiment, features of the desired channels can be filtered from the hyperspectral image to reconstruct the spectral channel reconstructed image. Alternatively, the dimensions of the subsequently extracted features can be aligned with the dimensions of the input image to facilitate subsequent feature enhancement of the input image based on the features of the desired channels.

[0010] In one possible implementation, the aforementioned process of filtering information from a preset spectral channel in a hyperspectral image to obtain a spectral channel reconstructed image includes: extracting features from the hyperspectral image to obtain a third feature; obtaining dynamic attention coefficients and spectral features based on the third feature using an attention mechanism; and fusing the dynamic attention coefficients, spectral features, and preset static attention coefficients to obtain a spectral channel reconstructed image.

[0011] In this embodiment of the application, when reconstructing channels of a hyperspectral image, an attention mechanism can be used, that is, the similarity between features can be used to reconstruct the required spectral features, which can make the features in the obtained channels more accurate and avoid data errors.

[0012] In one possible implementation, obtaining the second feature based on the features extracted from the spectral channel reconstructed image includes: downsampling the spectral channel reconstructed image at least once to obtain at least one downsampled spectral image; and extracting features from the at least one downsampled spectral image to obtain at least one scale feature, wherein the second feature includes at least one scale feature.

[0013] In this embodiment of the application, during the extraction of the second feature, one or more downsampling processes can be performed to obtain features at different scales. This allows for subsequent enhancement of features in the input image at different scales, enabling image enhancement or denoising at different scales, resulting in fused features with less noise, and thus an output image with less noise.

[0014] In one possible implementation, the aforementioned fusion of the first feature and the second feature to obtain the fused feature may further include: similar to the aforementioned downsampling of the hyperspectral channel reconstructed image, in order to align the various scales, the first feature may be downsampled at least once to obtain at least one downsampled feature, the scale of the at least one downsampled feature corresponding one-to-one with the scale of the at least one scale feature; and the at least one downsampled feature is fused with the corresponding scale feature to obtain at least one scale fused feature, and then the at least one scale fused feature is fused to obtain the fused feature. Therefore, in this embodiment, the first feature and the second feature can be fused at different scales, thereby achieving enhancement or denoising effects at each scale and reducing noise in the obtained fused feature.

[0015] In one possible implementation, the aforementioned method further includes: performing aberration estimation based on the hyperspectral image to obtain an aberration estimation map, the aberration estimation map being used to represent the error of pixels in the input image; and recovering the fusion features based on the aberration estimation map to obtain an output image.

[0016] In this embodiment, for possible aberrations, aberration estimation can be performed based on hyperspectral images to facilitate subsequent correction of possible aberrations in the input image and improve the accuracy of the obtained output image.

[0017] In one possible implementation, the aforementioned aberration estimation of the input image based on the hyperspectral image to obtain an aberration estimation map includes: using the hyperspectral image as input to a pre-trained generator and outputting an aberration estimation map, i.e., using the generator to perform aberration estimation online; or, using an offline calibration algorithm to perform aberration estimation of the input image based on the hyperspectral image to obtain an aberration estimation map, wherein the offline calibration algorithm is an algorithm determined in an offline state according to pre-set steps, i.e., using an offline algorithm to perform aberration estimation.

[0018] In this embodiment of the application, accurate aberration estimation can be achieved offline or online, so as to facilitate subsequent aberration correction.

[0019] In one possible implementation, the aforementioned fusion of the first feature and the second feature to obtain the fused feature may further include: aligning the first feature and the second feature, and merging the aligned first feature and the second feature to obtain the fused feature.

[0020] In this embodiment of the application, the first feature and the second feature can be aligned in spatial dimensions to improve the accuracy of fusing the first feature and the second feature and obtain a more accurate fused feature.

[0021] Secondly, this application provides an image processing apparatus, comprising:

[0022] The first input module is used to acquire the input image, which is the image captured by the image sensor.

[0023] The second input module is used to acquire hyperspectral images, the spectral channels of which are not exactly the same as those of the input image;

[0024] The first feature extraction module is used to extract features from the input image to obtain the first feature;

[0025] The second feature extraction module is used to obtain the second feature based on the features extracted from the hyperspectral image;

[0026] The feature fusion module is used to fuse the first feature and the second feature to obtain the fused feature.

[0027] The recovery module is used to obtain the output image based on the fusion features.

[0028] The effects achieved by the second aspect and any optional implementation of the second aspect can be referred to the description of the first aspect or any optional implementation of the first aspect, and will not be repeated here.

[0029] In one possible implementation, the feature fusion module is specifically used to fuse the first feature and the second feature based on an attention mechanism to obtain a fused feature.

[0030] In one possible implementation, the feature fusion module is further configured to: acquire a historical frame image, wherein the historical frame image is the previous frame image of the input image; extract features from the historical frame to obtain a third feature; and fuse the first feature, the second feature, and the third feature to obtain a fused feature.

[0031] In one possible implementation, the second feature extraction module is specifically used for: filtering spectral channel information from the hyperspectral image to obtain a spectral channel reconstructed image; and obtaining a second feature based on the features extracted from the spectral channel reconstructed image.

[0032] In one possible implementation, the second feature extraction module is specifically used for: extracting features from the hyperspectral image to obtain a third feature; obtaining dynamic attention coefficients and spectral features based on the third feature using an attention mechanism; and fusing the dynamic attention coefficients, spectral features, and preset static attention coefficients to obtain a spectral channel reconstructed image.

[0033] In one possible implementation, the second feature extraction module is specifically used to: downsample the spectral channel reconstructed image at least once to obtain at least one frame of downsampled spectral image; extract features from the at least one frame of downsampled spectral image to obtain at least one scale feature, wherein the second feature includes at least one scale feature.

[0034] In one possible implementation, the feature fusion module is further configured to: downsample the first feature at least once to obtain at least one downsampled feature, wherein the scale of the at least one downsampled feature corresponds one-to-one with the scale of the at least one scale feature; fuse the at least one downsampled feature with the corresponding scale feature to obtain at least one scale fused feature; and fuse the at least one scale fused feature to obtain a fused feature.

[0035] In one possible implementation, the aforementioned image processing apparatus further includes:

[0036] The aberration estimation module is used to perform aberration estimation based on the hyperspectral image and obtain an aberration estimation map, which is used to represent the error of the pixels in the input image.

[0037] The recovery module is specifically used to recover the fused features based on the aberration estimation map to obtain the output image.

[0038] In one possible implementation, the aberration estimation module is specifically used to: take the hyperspectral image as input to a pre-trained generator and output an aberration estimation map; or, use an offline calibration algorithm to perform aberration estimation on the input image based on the hyperspectral image to obtain an aberration estimation map, wherein the offline calibration algorithm is an algorithm determined in an offline state according to pre-set steps.

[0039] In one possible implementation, the feature fusion module is further configured to: align the first feature and the second feature, and fuse the aligned first feature and the second feature to obtain a fused feature.

[0040] Thirdly, embodiments of this application provide an image processing apparatus, including a processor and a memory, wherein the processor and the memory are interconnected via a circuit, and the processor calls program code in the memory to execute processing-related functions in the image processing method shown in any of the first aspects above. Optionally, the electronic device may be a chip.

[0041] Fourthly, embodiments of this application provide an electronic device, which may also be referred to as a digital processing chip or a chip. The chip includes a processing unit and a communication interface. The processing unit obtains program instructions through the communication interface, and the program instructions are executed by the processing unit. The processing unit is used to perform processing-related functions as described in the first aspect or any optional embodiment of the first aspect.

[0042] Fifthly, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the method described in the first aspect or any optional embodiment of the first aspect.

[0043] In a sixth aspect, embodiments of this application provide a computer program product containing instructions that, when run on a computer, cause the computer to perform the method described in the first aspect or any optional implementation thereof. Attached Figure Description

[0044] Figure 1 A schematic diagram of the structure of an electronic device provided in this application;

[0045] Figure 2 A schematic diagram of a system architecture provided for this application;

[0046] Figure 3 A schematic diagram comparing the method provided in this application with existing solutions;

[0047] Figure 4 A flowchart illustrating an image processing method provided in this application;

[0048] Figure 5 A schematic diagram of the structure of an image processing model provided in this application;

[0049] Figure 6 A schematic diagram of another image processing model provided in this application;

[0050] Figure 7 A schematic diagram of another image processing model provided in this application;

[0051] Figure 8 A schematic diagram of another image processing model provided in this application;

[0052] Figure 9 A schematic diagram of another image processing model provided in this application;

[0053] Figure 10 A schematic diagram of another image processing model provided in this application;

[0054] Figure 11A schematic diagram of another image processing model provided in this application;

[0055] Figure 12 A schematic diagram of another image processing model provided in this application;

[0056] Figure 13 A schematic diagram of another image processing model provided in this application;

[0057] Figure 14 A schematic diagram of another image processing model provided in this application;

[0058] Figure 15 A schematic diagram of another image processing model provided in this application;

[0059] Figure 16 A schematic diagram of another image processing model provided in this application;

[0060] Figure 17 A schematic diagram of another image processing model provided in this application;

[0061] Figure 18 A schematic diagram of the structure of an image processing device provided in this application;

[0062] Figure 19 A schematic diagram of another image processing apparatus provided in this application. Detailed Implementation

[0063] The technical solutions of the embodiments of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0064] First, the embodiments of this application involve a large number of applications related to neural networks. In order to better understand the solutions of the embodiments of this application, the relevant terms and concepts of neural networks that may be involved in the embodiments of this application will be introduced below.

[0065] (1) Neural Network

[0066] Neural networks can be composed of neural units, which can refer to units represented by x. s The arithmetic unit that takes input data and an intercept of 1 as input can output the following:

[0067]

[0068] Where s = 1, 2, ..., n, n is a natural number greater than 1, W s For x s The weight parameters are denoted by b, which represents the bias of the neural unit. f is the activation function of the neural unit, used to introduce non-linear characteristics into the neural network to convert the input signal into the output signal. The output signal of this activation function can be used as the input to the next convolutional layer; the activation function can be the sigmoid function. A neural network is a network formed by connecting multiple individual neural units, meaning the output of one neural unit can be the input of another. The input of each neural unit can be connected to the local receptive field of the previous layer to extract features from the local receptive field, which can be a region composed of several neural units.

[0069] (2) Loss Function

[0070] In training deep neural networks, to ensure the output closely approximates the desired predicted value, we compare the network's prediction with the target value and update the weight vector of each layer based on the difference. (Of course, there's usually a pre-configuration process before the first update, where parameters are pre-configured for each layer.) For example, if the prediction is too high, the weight vector is adjusted to predict a lower value. This adjustment continues until the deep neural network can predict the target value or a value very close to it. Therefore, we need to predefine "how to compare the difference between the predicted and target values," which is the loss function or objective function. These are important equations used to measure the difference between the predicted and target values. Taking the loss function as an example, a higher output value (loss) indicates a greater difference, so training the deep neural network becomes a process of minimizing this loss. Common loss functions include mean squared error, cross-entropy, logarithmic, and exponential loss functions. For example, mean squared error can be used as the loss function, defined as... The specific loss function can be selected based on the actual application scenario.

[0071] (3) Backpropagation algorithm

[0072] An algorithm for calculating the gradient of model parameters based on a loss function and updating the model parameters. Neural networks can use backpropagation (BP) to correct the initial parameter values ​​during training, thus reducing the reconstruction error loss. Specifically, forward propagation of the input signal to the output generates error loss; this error loss information is then propagated back to update the parameters of the initial neural network model, thereby converging the error loss. The backpropagation algorithm is an error-loss-driven backpropagation process aimed at obtaining the optimal parameters of the neural network model, such as the weight matrix.

[0073] In the embodiments of this application, the BP algorithm can be used to train the model during the training phase to obtain the trained model.

[0074] (4) Attention (also known as attention mechanism)

[0075] Attention mechanisms can quickly extract important features from sparse data. Attention occurs between the encoder and decoder, or more specifically, between the input and generated sentences. In contrast, the self-attention mechanism in a self-attention model occurs within the encoding matrix or the output sequence, extracting connections between distant words within the same sentence, such as syntactic features (phrase structure). Self-attention provides an effective modeling method for capturing global contextual information through QKV (key-value pairs). Assuming the input is Q (query), and the context is stored as key-value pairs (K, V), then the attention mechanism is essentially a mapping function from the query to a series of key-value pairs (key, value). The essence of the attention function can be described as a mapping from a query to a series of (key-value) pairs. Attention essentially assigns a weight coefficient to each element in the sequence, which can also be understood as soft addressing. If each element in the sequence is stored in (K, V) form, then attention performs addressing by calculating the similarity between Q and K. The similarity calculated between Q and K reflects the importance of the extracted V values, i.e., the weights, and then a weighted sum is obtained to obtain the final feature value.

[0076] Attention calculation mainly consists of three steps. The first step is to calculate the similarity between the query and each key to obtain weights. Common similarity functions include dot product, concatenation, and perceptron. The second step typically uses a softmax function (which can normalize the weights, resulting in a probability distribution where the sum of all weight coefficients is 1, and also highlights the weights of important elements) to normalize these weights. Finally, the weights and their corresponding key values ​​are weighted and summed to obtain the final feature value. The specific calculation formula is as follows:

[0077]

[0078] Where d is the dimension of matrix QK.

[0079] Furthermore, attention includes self-attention and cross-attention. Self-attention can be understood as a special type of attention where the inputs to the QKV features are consistent. Cross-attention, on the other hand, involves inconsistent inputs to the QKV features. Attention integrates the queried features as updated values ​​for the current features using the similarity between features (e.g., inner product) as weights. Self-attention is attention extracted based on the attention drawn from the feature map itself.

[0080] For convolutional networks, the kernel size limits the receptive field, often requiring multiple layers to focus on the entire feature map. Self-attention, on the other hand, offers the advantage of global focus; it can acquire global spatial information of the feature map through simple queries and assignments. A unique aspect of self-attention in query-key-value (QKV) models is that the inputs for each QKV value are consistent.

[0081] (5) Raw data: Raw data records the original information of the camera sensor. It is an unprocessed and uncompressed format. RAW can be conceptualized as "raw image encoded data" or more figuratively called "digital negative".

[0082] (6)R(red), G(green), B(blue)

[0083] In this image, R represents red, G represents green, and B represents blue. Each image can be represented by the color values ​​of these three channels. For example, an RGB image represents an image with three color channels, and an RGGB image represents an image with four color channels, two of which are G channels.

[0084] (7) Hyperspectral image

[0085] Hyperspectral sensors are used to image the ultraviolet, visible, near-infrared, or mid-infrared regions of the electromagnetic spectrum. The number of channels in a hyperspectral image is typically determined by the number of spectral bands acquired by the hyperspectral sensor.

[0086] The method provided in this application can be applied to various image processing scenarios, such as image denoising, image enhancement, or image super-resolution.

[0087] Specifically, regarding the deployment methods provided in this application, there are several deployment methods. For example, the method provided in this application can be deployed in electronic devices, allowing users to directly perform image processing on the electronic devices; alternatively, it can be deployed in a cloud platform to provide image processing services to user terminals. The different deployment methods are described below.

[0088] Deployment Method 1: Deployed in electronic devices

[0089] The electronic devices provided in this application embodiment may specifically include handheld devices, in-vehicle devices, data processing devices, computing devices, and other electronic devices including or connected to image sensors. They may also include digital cameras, cellular phones, cameras, smartphones, personal digital assistant (PDA) computers, tablet computers, laptop computers, machine type communication (MTC) terminals, point of sale (POS) terminals, in-vehicle computers, head-mounted devices, data processing devices (such as wristbands, smartwatches, etc.), security equipment, virtual reality (VR) devices, augmented reality (AR) devices, and other electronic devices with imaging capabilities.

[0090] Taking digital cameras as an example, a digital camera, short for digital camera, is a type of camera that uses photoelectric sensors to convert optical images into digital signals. Unlike traditional cameras that rely on changes in photosensitive chemicals on film to record images, digital camera sensors are photosensitive charge-coupled devices (CCDs) or complementary metal-oxide-semiconductor (CMOS) sensors. Compared to traditional cameras, digital cameras, by directly using photoelectric conversion image sensors, offer advantages such as greater convenience, speed, repeatability, and timeliness. With the development of CMOS processing technology, digital cameras have become increasingly powerful, almost completely replacing traditional film cameras, and are widely used in consumer electronics, human-computer interaction, computer vision, autonomous driving, and other fields.

[0091] For example, Figure 1 A schematic diagram of an electronic device provided in this application is shown. As shown, the electronic device may include a lens assembly 110, an image sensor 120, and an electrical signal processor 130. The electrical signal processor 130 may include an analog-to-digital (A / D) converter 131 and a digital signal processor 132. The analog-to-digital converter 131 is an analog-to-digital signal converter used to convert analog electrical signals into digital electrical signals.

[0092] It should be understood that Figure 1 The electronic devices shown are not limited to those mentioned above, and may include more or fewer other devices, such as batteries, flashlights, buttons, sensors, etc. This application embodiment only uses an electronic device equipped with an image sensor 120 as an example for illustration, but the components installed on the electronic device are not limited to this.

[0093] In this application embodiment, the aforementioned sensor 120 may specifically include an image sensor, a multispectral image (MSI) sensor, etc.

[0094] The light signal reflected from the object is converged by the lens assembly 110 and imaged onto the sensor 120. The image sensor 120 converts the light signal into an analog electrical signal. The analog electrical signal is converted into a digital electrical signal by the analog-to-digital (A / D) converter 131 in the electrical signal processor 130, and then processed by the digital signal processor 132, for example, by optimizing the data electrical signal through a series of complex mathematical algorithms, and finally outputting an image. The electrical signal processor 130 may also include an analog signal preprocessor 133, which is used to preprocess the analog electrical signal transmitted from the image sensor before outputting it to the analog-to-digital converter 131.

[0095] The performance of an image sensor affects the quality of the final output image. An image sensor, also known as a photosensitive chip or photosensitive element, contains hundreds of thousands to millions of photoelectric conversion elements. When exposed to light, these elements generate electrical charges, which are then converted into digital signals by an analog-to-digital converter chip. An image sensor essentially comprises multiple pixels of photosensitive elements, achieving image formation through photoelectric response.

[0096] The MSI sensor can simultaneously acquire image signals across multiple spectral bands, including both visible and invisible spectral bands (such as infrared and ultraviolet bands), and has the potential to improve noise reduction in RGB images and videos. In some scenarios (such as extremely low-light scenes), RGB images can no longer capture the details and texture information of the scene being captured, while MSI images can cover a wider spectral range, capturing more detailed information in the non-visible spectral bands. Therefore, the method provided in this application can be specifically deployed in the electrical signal processor 130 of an electronic device, such as in the digital signal processor 132, or in other processors of the electronic device. Specifically, the method provided in this application can utilize hyperspectral images to denoise visible light images, that is, to use the richer information contained in the hyperspectral image to filter out noise in the visible light image, obtaining a denoised visible light image.

[0097] Meanwhile, MSI images can cover more spectral bands within the visible spectrum. Under the same exposure conditions, MSI sensors receive more photons, resulting in images with a higher signal-to-noise ratio (SNR). Using MSI images with higher SNR as additional input to assist and guide the RGB image noise reduction process can effectively improve noise reduction results and restore more texture details.

[0098] Deployment Method 2: Deployed on a cloud platform

[0099] This application also provides a cloud platform through which one or more terminals can access. The method provided in this application can be deployed in a cloud platform to provide image noise reduction or enhancement services to terminals.

[0100] For example, Figure 2 One application scenario for the method provided in this application includes a cloud platform 11 and a terminal 12. The cloud platform 11 and the terminal 12 can be connected via wired or wireless means.

[0101] The cloud platform 11 may specifically include a server cluster with storage and processing functions. The method provided in this application embodiment can be deployed in the cloud platform 11, specifically receiving multiple frames of images from the terminal 12, performing image processing based on the multiple frames of images, such as image denoising and enhancement, and feeding back the processed image to the terminal 12.

[0102] Terminal 12 can perform image processing by interacting with cloud platform 11. This terminal can specifically include, but is not limited to, personal computers, computer workstations, smartphones, tablets, laptops, and smart cars. Terminal 12 can transmit images to cloud platform 11. These images can be images captured by the terminal itself, images input by the user, or images stored locally on the terminal. For example, the cloud platform can provide services to users through clients deployed on the terminal or web pages on the terminal. Taking the deployment of a client on the terminal as an example, the user can send images captured by the terminal to the cloud platform through the client deployed on the terminal, such as transmitting RGB and MSI images. Cloud platform 11, using the method provided in this application, denoises or enhances the RGB image, outputs a high-definition RGB image, and feeds it back to terminal 12.

[0103] In one possible scenario, it can also be applied to a scenario with multiple terminals. For example, a user can use a terminal other than the aforementioned terminal 12 to capture the aforementioned RGB image and MSI image, transmit the RGB image and MSI image to the aforementioned terminal 12, and the terminal 12 can upload the RGB image and MSI image to the cloud platform 11 for image denoising or enhancement processing, and then send the enhanced image back to the terminal 12. The terminal 12 can then send the enhanced image back to the terminal that captured the aforementioned RGB image and MSI image.

[0104] Typically, the imaging process of digital image sensors is affected by noise, such as noise from the imaging module and circuitry. Noise usually manifests in images as specular noise, blur, or artifacts, leading to a degrade in image quality. For example, in some scenarios, such as low-light or high-light scenes, there may be a loss of image details or textures.

[0105] For example, one existing approach utilizes RGB-assisted MSI texture detail enhancement, taking a low-quality MSI image and a high-quality RGB image as input, and outputting a high-quality MSI image. This type of approach primarily leverages the higher spatial resolution of RGB images and videos to improve the enhancement and restoration of MSI images. However, this approach does not yield a denoised RGB image.

[0106] For example, one existing approach utilizes NIR / UV-assisted RGB texture detail enhancement. In this type of approach, the input consists of a single non-visible light band image / video and an RGB image / video, with the output being a high-quality RGB image / video. This approach processes the NIR image and obtains the fusion weights for the RGB images, then fuses the RGB images according to these weights, ultimately improving the quality of the output RGB image. However, this approach is only applicable to specific scenarios and has requirements regarding the degradation mode of the input image, such as extremely low-light imaging or foggy / hazy scene imaging. Furthermore, this type of approach only introduces the non-visible light band and does not utilize the full-band information of MSI, resulting in limited improvement.

[0107] For example, in an existing scheme, it was found that green channel data has a higher signal-to-noise ratio and more detailed texture information. The input is a low-quality RGB image and a green channel image, and the output is a high-quality RGB image. Therefore, this scheme uses green spectral channel data to extract green channel guiding features and fuses these features with features from the RGB image reconstruction network to obtain the final output. However, using only a small amount of spectral information as guiding information results in limited performance improvement. Furthermore, the spectral information is spatially aligned with the RGB image by default, making it unsuitable for multi-camera imaging systems.

[0108] Therefore, the image processing method provided in this application can utilize the characteristics of multispectral data to assist in image restoration and improve image quality.

[0109] A comparison with the aforementioned existing solutions can be made as follows: Figure 3 As shown, the first type of scheme utilizes RGB-assisted MSI texture detail enhancement, while the second and third types utilize RGB-assisted MSI texture detail enhancement and green spectral channel data to enhance RGB images, respectively. Clearly, compared to the first type of scheme, the method provided in this application takes RGB and hyperspectral images as input and outputs a high-quality RGB image. Its input and output are different from those of the first, second, and third types of schemes.

[0110] The method and process provided in this application are described below.

[0111] See Figure 4 The present application provides a schematic flowchart of an image processing method.

[0112] 401. Obtain the input image.

[0113] The input image may specifically include an image captured using the aforementioned electronic device, or an image read from storage, etc.

[0114] For example, if the method provided in this application embodiment is deployed in an electronic device, the input image can be an image captured by the electronic device, which can refer to the foregoing. Figure 1 The corresponding descriptions will not be repeated here.

[0115] For example, if the method provided in this application is deployed in a cloud platform, the input image can specifically be an image received from a terminal. For instance, a user can input an image on the terminal, or take a picture using the terminal and send the image to the cloud platform via the terminal.

[0116] Specifically, the input image can include an image with at least one channel. For example, it can include a raw image or an RGB image, to accommodate various visible light denoising scenarios.

[0117] Optionally, during the process of acquiring the input image, after acquiring the image collected by the image sensor, preprocessing operations such as image alignment or color space conversion can be performed to obtain an image that better conforms to the preset specifications, so as to facilitate further processing of the input image.

[0118] 402. Acquire hyperspectral images.

[0119] The hyperspectral image (HSI) can specifically be an image obtained using data acquired by a hyperspectral sensor, and the hyperspectral image can include image information from multiple channels.

[0120] The spectral channels of the hyperspectral image are not exactly the same as those of the input image. For example, the number of channels in the hyperspectral image can be greater than the number of channels in the input image, or the hyperspectral image may include information about channels not included in the input image.

[0121] Furthermore, the acquisition time of the input image is the same as or the difference between the acquisition time of the hyperspectral image is less than a preset value, thereby ensuring that the environmental information contained in the hyperspectral image and the input image is the same or similar, so that the texture and details of the input image can be recovered by using the hyperspectral image later.

[0122] Optionally, during the acquisition of HSI, after acquiring the image collected by the HSI sensor, preprocessing operations such as image alignment or color space conversion can be performed to obtain an image that better conforms to the preset specifications, so as to facilitate further processing of HSI.

[0123] It should be noted that this application does not limit the execution order of steps 401 and 402. Step 401 can be executed first, or step 402 can be executed first, or steps 401 and 402 can be executed simultaneously. The specific order can be determined according to the actual application scenario.

[0124] 403. Extract features from the input image to obtain the first feature.

[0125] After obtaining the input image, features can be extracted from the input image using a neural network. For ease of differentiation, the features extracted from the input image are called the first features.

[0126] In one possible implementation, if it is necessary to fuse features at multiple scales in the subsequent fusion step, multiple scale features can be extracted from the first feature when extracting the first feature; or, when extracting the first feature, the input image can be downsampled once or multiple times to obtain images at different scales, and features can be extracted from the images at different scales to obtain features at different scales, which are the first features.

[0127] 404. Based on the features extracted from the hyperspectral image, the second feature is obtained.

[0128] After obtaining a hyperspectral image, features can be extracted from the hyperspectral image using a neural network. For ease of differentiation, the features extracted from the hyperspectral image are called the second features.

[0129] In one possible implementation, information of the desired spectral channels can be filtered from a hyperspectral image to obtain a reconstructed spectral channel image; subsequently, a second feature is obtained based on the features extracted from the reconstructed spectral channel image. Therefore, in this embodiment, information of the actual desired spectral channels can be filtered from a hyperspectral image, and a reconstructed spectral channel image can be obtained to extract the features of the desired spectral channels.

[0130] Furthermore, in one possible implementation, features can be extracted from the hyperspectral image to obtain a third feature; dynamic attention coefficients and spectral features can be obtained based on the third feature using an attention mechanism; and the dynamic attention coefficients, spectral features, and preset static attention coefficients can be fused to obtain a reconstructed spectral channel image. Therefore, in this embodiment, the desired spectral channel can be reconstructed based on an attention mechanism to obtain the features of the desired spectral channel more accurately.

[0131] Furthermore, when extracting the second feature, features at multiple scales can be extracted to assist in denoising the input image from multiple scales. Specifically, the spectral channel reconstructed image can be downsampled at least once to obtain at least one downsampled spectral image; features can be extracted from the at least one downsampled spectral image to obtain at least one scale feature, wherein the second feature includes at least one scale feature.

[0132] 405. By fusing the first feature and the second feature, a fused feature is obtained.

[0133] After extracting the first and second features, they can be fused to obtain a fused feature. This can be achieved through weighted fusion or other fusion mechanisms, outputting a fused feature that combines the first and second features. This fused feature can be understood as a fusion of features extracted from the input image and hyperspectral features extracted from the hyperspectral image; that is, it carries the data distribution characteristics of the input image and the hyperspectral features of the hyperspectral image. Therefore, the hyperspectral features in the hyperspectral image can be used to supplement or enhance the features extracted from the input image in dimensions such as texture or detail, resulting in a fused feature containing richer features.

[0134] Optionally, the first feature is downsampled at least once to obtain at least one downsampled feature, the scale of which corresponds one-to-one with the scale of the at least one scale feature; the at least one downsampled feature is fused with its corresponding scale feature to obtain at least one scale fused feature; the at least one scale fused feature is then fused to obtain a fused feature. In this embodiment, feature fusion can be performed at multiple different scales to obtain features at different scales. Therefore, spectral features at different scales extracted from hyperspectral images can be used to smooth features extracted from visible light images to achieve denoising effects at different scales.

[0135] 406. Obtain the output image based on the fusion features.

[0136] After obtaining the fusion features, the image can be restored using these features to obtain the final output image, which is the output image after denoising the input image.

[0137] In this embodiment, features can be extracted from both the input image and the hyperspectral image to be denoised. Hyperspectral features that represent richer details and textures can be extracted from the hyperspectral image. These features are then fused with the features extracted from the input image to obtain fused features, thereby achieving a denoising effect on the input image. Therefore, in this embodiment, less noisy features extracted from the hyperspectral image can be used to fuse the features extracted from the input image, thus smoothing the noise in the input image's features. This results in less noisy fused features, leading to a less noisy output image after image restoration, achieving a denoising effect on the input image.

[0138] In one possible implementation, the first feature and the second feature are aligned before fusing the first feature and the second feature, and then the aligned first feature and the second feature can be fused to obtain the fused feature.

[0139] In one alternative implementation, to correct aberrations in the input image, aberration estimation can be performed based on the hyperspectral image, and the aberration estimation map can be used to correct the aberrations in the input image. That is, image restoration can be performed based on the aberration estimation map and fused features, thereby using the aberration estimation map to correct the aberrations of individual pixels in the image, resulting in a more accurate output image.

[0140] Optionally, the hyperspectral image can be used as input to a pre-trained generator to output an aberration estimation map; or, an offline calibration algorithm can be used to perform aberration estimation on the input image based on the hyperspectral image to obtain an aberration estimation map. The offline calibration algorithm is an algorithm determined offline according to pre-set steps. In this embodiment of the application, aberration estimation of the input image can be achieved in various ways to achieve more accurate aberration estimation.

[0141] The foregoing has provided a combined introduction to the method flow provided in this application. The following section will introduce the method flow provided in this application in conjunction with a specific image processing model.

[0142] For example, this application provides an image and video restoration system for a multimodal camera, which receives input data from an HSI sensor and an RGB sensor. The RGB sensor can be replaced with other visible light sensors, and the HSI sensor can be replaced with other sensors that acquire multispectral images; the specific method can be determined according to the actual application scenario. Low-quality RGB images or videos are obtained from the RGB camera, and the HSI image or video is used to guide the RGB image or video restoration process, performing functions such as noise reduction and detail recovery.

[0143] Typically, for single-frame images and videos, combined with the aforementioned... Figure 4 The corresponding method flow can be configured with different modules.

[0144] For example, for single-frame image processing, see [link / reference]. Figure 5 This application provides a schematic diagram of the structure of an image processing model.

[0145] The image processing model may include, but is not limited to: HSI preprocessing module 100, band filtering and fusion module 201, multi-scale feature extraction module 202, multi-scale attention fusion module 300, aberration estimation module 400, RGB preprocessing module 500 and restoration backbone network 600, etc.

[0146] in:

[0147] HSI preprocessing module 100: Used to receive input HIS images. C in >3, C inIndicates the number of spectral channels in the HIS; performs image alignment, color space conversion, and other processing operations, and outputs the processed image. For monocular imaging systems, spatial alignment and color space conversion of images are not required, so this module is not needed or does not need to be enabled.

[0148] Spectral band filtering and fusion module 201: Input is an HSI image, output is feature tensor data. Specifically, it can receive an input image or an image output from the HSI preprocessing module 100. The spectral band selection and fusion module contains a multi-layer convolutional structure. This structure is used for feature extraction and spectral selection to reconstruct spectral channels, and the output is... That is, the aforementioned spectral channel reconstructed image.

[0149] Multi-scale feature extraction module 202: Input the filtered HSI image and small-scale HSI images after downsampling. The spectral multi-scale guided feature extraction module, i = 2, 4, ..., N, contains a multi-layer convolutional structure to perform deep feature extraction on multi-scale images, improving the spectral feature representation capability and obtaining a multi-scale feature tensor. i = 2, 4, ..., N, which is the aforementioned second feature, serves as the input to the subsequent fusion module.

[0150] Multi-scale attention fusion module 300: A feature space fusion module based on multi-scale attention. The inputs are the feature tensors obtained from the multi-scale feature extraction module 202 and the RAW / RGB image. The output is the fused multiple feature tensors. Specifically, the input is the multi-scale features of module 202. and the image of module 500 This module consists of multi-layer convolutional layers, a transformer module, fully connected layers, upsampling layers, and downsampling layers. The multi-layer convolutional layers are used to extract images. The features in the image, namely the first feature mentioned above, are then processed through a transformer module, a fully connected layer, upsampling and downsampling layers to align and fuse image features and spectral features, resulting in spatially consistent fused features with a higher signal-to-noise ratio. The aforementioned fusion features are used for processing in subsequent modules.

[0151] Aberration estimation module 400: Input is an HSI image, output is an optical aberration estimation map. Receives input HSI image: The optical aberration estimation module consists of a point spread function generator, which performs online estimation of optical aberrations or directly uses offline calibration to obtain an optical aberration estimation map.

[0152] RGB preprocessing module 500: Receives input RGB images. C v ∈[1,3], performs image alignment, color space conversion and other processing modules, and outputs the processed image. Monocular systems do not require spatial alignment and color space conversion of images, so this module is not needed.

[0153] Denoising and detail restoration backbone network 600: Input is the feature Z of module 300 and the image of module 500. The noise reduction and detail restoration backbone network (also known as the restoration backbone network) contains multiple convolutional layers for feature extraction and image reconstruction, outputting I... out ∈ C out ∈[1,3], which is the aforementioned output image.

[0154] For example, regarding video image processing, see [reference needed]. Figure 6 This application provides a schematic diagram of the architecture of an image processing model.

[0155] As mentioned above Figure 5 The difference lies in the fact that, for RGB image input, a processing channel for historical frames is also set up. For example... Figure 6 As shown, when denoising the RGB image of frame T, the RGB image of frame T-1 is also introduced, and features are extracted from the RGB image of frame T and the RGB image of frame T-1 respectively, as follows: Figure 6 The feature extraction module 601 shown can fuse the features extracted from the RGB image of frame T and the RGB image of frame T-1 in the multi-scale attention fusion module 300, so as to combine the information in the historical frames to further assist the current frame in denoising and improve the image denoising effect.

[0156] The steps performed by each module can be summarized in the foregoing. Figure 4 The corresponding methodology and workflow are described below, with each major module's specific functions and structure explained in detail.

[0157] 1. Band filtering and fusion module 201

[0158] The specific steps performed by the band selection and fusion module 201 can be found in the aforementioned document. Figure 4 The description of the reconstructed spectral channel output spectral channel reconstructed image in step 404.

[0159] For example, the structure of the band selection and fusion module 201 can be as follows: Figure 7As shown, the band filtering and fusion module 201 may include convolutional layers and one or more spectral channel attention blocks (SCABs). The convolutional layers can be used to extract features, and the SCABs can be used to further extract the desired spectral features based on the attention mechanism. The input to the band filtering and fusion module 201 is... The output is

[0160] Specifically, the structure of SCAB can be as follows: Figure 8 As shown, this includes convolutional layers and multi-layer perceptron (MLP) layers. The input to SCAB can be represented as a feature tensor. After passing through a convolutional layer and an MLP layer, the dynamic attention coefficients W1 (with dimensions 1×1×C) are obtained. f1 This refers to the attention coefficient matrix dynamically calculated during image processing, which is then fused with the input features to obtain the dynamically fused spectral features.

[0161] x f2 (i,j)=x f1 (i,j)×W1

[0162] The dynamically fused features yield fused spectral features, which are then processed by an activation function (here, the activation function is the Gaussian error linear unit, GELU) and combined with the static attention parameters W2 (of dimension C) obtained during the training phase. f2 ×C f1 The spectral characteristics after static fusion were obtained:

[0163] x i+1 (i,j)=x f2 (i,j)×W2

[0164] This outputs the fused spectral characteristics.

[0165] In this embodiment, features can be extracted from the input image based on an attention mechanism, and the desired spectral features can be extracted based on dynamic attention and static attention respectively, thereby selecting the spectral features that are used to assist in the subsequent RGB image denoising.

[0166] 2. Multi-scale feature extraction module 202

[0167] The specific steps performed by the multi-scale feature extraction module 202 can be found in the description of feature extraction from the hyperspectral image in step 404 above.

[0168] For example, the structure of the multi-scale feature extraction module 202 can be as follows: Figure 9As shown, the multi-scale feature extraction module 202 may include processing channels of multiple scales. Each channel may include a downsampling operation and a feature extraction module. The downsampling operation can be used to downsample the input image, and the feature extraction module can be used to extract features.

[0169] For example, the input to the multi-scale feature extraction module 202 can be represented as: Obtained through the feature extraction module In other channels, By downsampling, low-resolution input features are obtained. Obtained through the feature extraction module

[0170] In this embodiment, features at different scales can be extracted, so that when performing feature fusion in the future, the noise of the RGB image can be reduced from different scales, and the noise in the RGB image can be further smoothed by using HIS features at different scales, thereby improving the SNR of the output image.

[0171] 3. Multi-scale attention fusion module 300

[0172] The steps performed by the multi-scale attention fusion module 300 can be found in the description of step 405 mentioned above; similar steps will not be repeated here.

[0173] Specifically, the multi-scale attention fusion module 300 can be used to achieve simultaneous fusion of spatial, temporal, and spectral data. For example, the structure of the multi-scale attention fusion module 300 can be as follows: Figure 10 As shown, in this embodiment, feature extraction and fusion can be performed at multiple scales. A downsampling module, convolutional layers, and upsampling modules can be set separately. The downsampling module is used to downsample the input data, the convolutional layers are used for feature extraction, and a spatial-temporal-temporal fusion module (Spectral-Spatial-Temporal Transformer Block, STB) is used. For the input RGB image (i.e.... Figure 10 The current frame shown Historical Frames Features can be extracted from different scales, and STB can be used to fuse the extracted spectral features of the same scale. This allows for the fusion of RGB features (i.e. features extracted from RGB images) and spectral features at different scales, and the fused low-scale features can be upsampled to restore the scale.

[0174] Specifically, the network structure of STB differs for different scenarios involving images and videos.

[0175] For example, in image restoration scenarios, taking one scale as an example, the STB (Structured Block) as a feature spectral spatial fusion module can be structured as follows: Figure 11 As shown, STB spatially divides features into non-overlapping feature blocks and performs a fusion operation on feature blocks at the same location. This is applicable to features at different scales. Linear projection is performed using an attention mechanism, and is divided into... And the input image Linear projection based on attention mechanisms is divided into... right and as well as and Perform inner product calculations separately, normalize the calculated inner products, and then concatenate them. Then, analyze the features of the concatenated product. and After performing the multiplication operation, the fused features are output.

[0176] For example, in video restoration scenarios, STP is extended to feature spectral, spatial, and temporal fusion modules, and the structure of STB can be as follows: Figure 12 As shown, a fusion operation is performed on different time feature blocks at the same location. Similar to the aforementioned... Figure 11 The difference lies in the addition of scale features extracted from historical frames. The integration process is about to begin. After linear projection, it is divided into right and When performing an inner product, during the multiplication operation, the number of steps increases. and inner product and This results in a fused feature that incorporates features from historical frames.

[0177] 4. Aberration Estimation Module 400

[0178] The aberration estimation module 400 can be used to implement the steps described in step 406 above, which involve aberration estimation and outputting an aberration estimation map.

[0179] Specifically, aberration estimation can be implemented in various ways, such as online or offline methods, which will be introduced below.

[0180] (1) Online distortion and point spread function estimation

[0181] For example, such as Figure 13 As shown, HSI can be used as input to a point spread function generator, which outputs an optical aberration estimate. The point spread function generator is pre-trained. The structure of this point spread function generator can be as follows: Figure 14As shown, the point spread function generator can include multiple convolutional layers.

[0182] (2) Offline distortion and point spread function estimation

[0183] An optical aberration estimation map is obtained using an offline calibration method. During the network inference stage, this offline-calibrated optical aberration estimation map is directly used as input to the denoising and detail restoration network. The offline calibration flowchart is shown below. Figure 15 As shown, after the image is acquired Then, spectral alignment, spatial alignment, and fuzzy kernel estimation are performed to output the result. During the online inference phase, the steps calibrated offline can be used directly to output the corresponding aberration estimation map.

[0184] 5. Restore backbone network 600

[0185] The recovery backbone network 600 can be used to perform the step 406 described above to obtain the output image. Specifically, a pre-trained recovery backbone network can be used to output the final denoised output image.

[0186] For example, such as Figure 16 As shown, features can be fused. The aberration estimation map is used as input to the pre-trained restoration backbone network, which performs the image restoration steps to obtain the output image.

[0187] This application provides an image processing architecture applicable to various image restoration tasks that use RAW / RGB images as input, such as noise reduction, depigmentation, and super-resolution. The modules within this architecture can also be adapted to other types of image restoration and enhancement tasks that require RAW / RGB images. Furthermore, the image processing solution provided in this application exhibits high robustness to real-world open scenes and is adaptable to RAW images acquired using different sensor models and in different imaging scenarios.

[0188] Accordingly, the loss value that needs to be calculated during the training phase can be as follows: Figure 17As shown, the L1 loss between the aberration estimation map and the ground truth aberration map, as well as the L1 loss between the output RGB image and the ground truth image, can be calculated. The model is then updated based on these loss values ​​to obtain the updated image processing model. This updated model can then be deployed to denoise the input image using HIS during the inference phase. Specifically, during the training phase, the L1 reconstruction error (L1 loss) and other general reconstruction losses are used to constrain the relationship between the network output and the ground truth image. Simultaneously, the L1 reconstruction error can be used to constrain the relationship between the optical aberration estimation map and the auxiliary ground truth image in the optical aberration estimation module 400.

[0189] L1 loss can be expressed as:

[0190]

[0191] Therefore, during the inference phase, the trained model can be used directly to perform inference tasks.

[0192] Therefore, the method provided in this application implements a hyperspectral-guided RGB image / video noise reduction and detail restoration framework, with input T v (T v ≥1) Frame RGB image and T h (T h ≥1) frames of full-spectrum hyperspectral image, output T o (T o ≥1) RGB images with denoising and detail enhancement. Light is used to extract guiding features from the spectral image, which are then used for subsequent RGB fusion and denoising. This allows for the fusion of RGB features using spectral features, enhancing the texture and detail of the RGB features while reducing noise. Furthermore, the spectral feature-RGB feature fusion and denoising are not independent algorithm modules but are directly performed by a parameter-optimizable network, outputting the denoised result.

[0193] Compared with existing solutions, such as those that utilize only a single or a few bands, which offer limited incremental information and thus limited performance improvement, the method provided in this application can utilize full-spectrum information, providing more texture information and higher signal-to-noise ratio auxiliary data, significantly improving RAW or RGB noise reduction and detail recovery.

[0194] For example, compared to fusing HSI and RGB in the spectral dimension, which requires higher spatial registration performance from HSI and RGB binocular systems, the solution provided in this application simultaneously fuses HSI and RGB images in the spectral, spatial, and temporal dimensions. The feature fusion module is uniformly optimized end-to-end with other modules in the embodiments of this application, improving the fusion effect and the final RGB image / video restoration effect.

[0195] Furthermore, the solution provided in this application offers optical aberration estimation and correction, which processes and restores blur and chromatic aberration caused by optical path distortion. Specifically, it can correct the aberrations of the imaging system online or offline, restoring more details.

[0196] The method flow provided in this application has been described above. The structure of the apparatus provided in this application for executing the aforementioned method flow is described below.

[0197] See Figure 18 This application provides a schematic diagram of an image processing apparatus, which includes:

[0198] The first input module 1801 is used to acquire an input image, which is an image captured by an image sensor.

[0199] The second input module 1802 is used to acquire a hyperspectral image, the spectral channels of which are not exactly the same as those of the input image;

[0200] The first feature extraction module 1803 is used to extract features from the input image to obtain the first feature;

[0201] The second feature extraction module 1804 is used to obtain a second feature based on the features extracted from the hyperspectral image;

[0202] The feature fusion module 1805 is used to fuse the first feature and the second feature to obtain the fused feature;

[0203] Recovery module 1806 is used to obtain the output image based on the fusion features.

[0204] In one possible implementation, the feature fusion module 1805 is specifically used to fuse the first feature and the second feature based on an attention mechanism to obtain a fused feature.

[0205] In one possible implementation, the feature fusion module 1805 is further configured to: acquire a historical frame image, wherein the historical frame image is the previous frame image of the input image; extract features from the historical frame to obtain a third feature; and fuse the first feature, the second feature and the third feature to obtain a fused feature.

[0206] In one possible implementation, the second feature extraction module 1804 is specifically used for: filtering information of spectral channels from the hyperspectral image to obtain a spectral channel reconstructed image; and obtaining a second feature based on the features extracted from the spectral channel reconstructed image.

[0207] In one possible implementation, the second feature extraction module 1804 is specifically used to: extract features from the hyperspectral image to obtain a third feature; obtain dynamic attention coefficients and spectral features based on the third feature using an attention mechanism; and fuse the dynamic attention coefficients, spectral features, and preset static attention coefficients to obtain a spectral channel reconstructed image.

[0208] In one possible implementation, the second feature extraction module 1804 is specifically used to: downsample the spectral channel reconstructed image at least once to obtain at least one frame of downsampled spectral image; extract features from the at least one frame of downsampled spectral image to obtain at least one scale feature, wherein the second feature includes at least one scale feature.

[0209] In one possible implementation, the feature fusion module 1805 is further configured to: downsample the first feature at least once to obtain at least one downsampled feature, wherein the scale of the at least one downsampled feature corresponds one-to-one with the scale of the at least one scale feature; fuse the at least one downsampled feature with the corresponding scale feature to obtain at least one scale fusion feature; and fuse the at least one scale fusion feature to obtain a fused feature.

[0210] In one possible implementation, the aforementioned image processing apparatus further includes:

[0211] The aberration estimation module 1807 is used to perform aberration estimation based on the hyperspectral image to obtain an aberration estimation map, which is used to represent the error of the pixels in the input image.

[0212] The recovery module 1806 is specifically used to recover the fused features based on the aberration estimation map to obtain the output image.

[0213] In one possible implementation, the aberration estimation module 1807 is specifically used to: take the hyperspectral image as input to a pre-trained generator and output an aberration estimation map; or, use an offline calibration algorithm to perform aberration estimation on the input image based on the hyperspectral image to obtain an aberration estimation map, wherein the offline calibration algorithm is an algorithm determined in an offline state according to a pre-set step.

[0214] In one possible implementation, the feature fusion module 1805 is further configured to: align the first feature and the second feature, and fuse the aligned first feature and the second feature to obtain a fused feature.

[0215] like Figure 19 The diagram shown is a hardware structure schematic of an image processing apparatus 190 provided in an embodiment of this application. This image processing apparatus 190 can be used to implement the aforementioned... Figures 4 to 17 The steps of the method in the text.

[0216] Figure 19 The image processing apparatus 190 shown may include a processor 1901, a memory 1902, a communication interface 1903, and a bus 1904. The processor 1901, the memory 1902, and the communication interface 1903 can be connected to each other via the bus 1904.

[0217] The processor 1901 is the control center of the image processing device 190. It can be a general-purpose central processing unit (CPU) or other general-purpose processors. The general-purpose processor can be a microprocessor or any conventional processor, such as a GPU or NPU, and can be adapted to the actual application scenario.

[0218] As an example, processor 1901 may include one or more CPUs, and may also include other processors, such as... Figure 19 The CPU, NPU, or GPU shown are examples of such devices.

[0219] The memory 1902 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto.

[0220] In one possible implementation, the memory 1902 can exist independently of the processor 1901. The memory 1902 can be connected to the processor 1901 via a bus 1904, and is used to store data, instructions, or program code. When the processor 1901 calls and executes the instructions or program code stored in the memory 1902, it can implement the methods provided in the embodiments of this application, for example, Figures 4 to 17 The method shown.

[0221] In another possible implementation, the memory 1902 can also be integrated with the processor 1901.

[0222] The communication interface 1903 is used for the image processing device 190 to connect with other devices via a communication network, which may be Ethernet, radio access network (RAN), wireless local area network (WLAN), etc. The communication interface 1903 may include a receiving unit for receiving data and a transmitting unit for transmitting data.

[0223] The 1904 bus can be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, or an extended industry standard architecture (EISA) bus. This bus can be divided into address bus, data bus, and control bus, etc. For ease of representation, Figure 19 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0224] It should be pointed out that, Figure 19 The structure shown does not constitute a limitation on the image processing device 190, except... Figure 19 In addition to the components shown, the image processing apparatus 190 may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0225] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc., including several instructions to cause a device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0226] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0227] This application also provides a computer-readable storage medium storing a program for training a model or performing inference tasks, which, when run on a computer, causes the computer to perform the aforementioned... Figures 3 to 8 All or part of the steps in the method described in the embodiments shown.

[0228] This application also provides a digital processing chip. This digital processing chip integrates circuitry for implementing the aforementioned processor or processor functions, and one or more interfaces. When the digital processing chip integrates a memory, it can perform the method steps of any one or more of the foregoing embodiments. When the digital processing chip does not integrate a memory, it can be connected to an external memory via a communication interface. The digital processing chip implements the method steps of any one or more of the foregoing embodiments based on the program code stored in the external memory.

[0229] This application also provides a computer program product comprising one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., a solid-state disk (SSD)).

[0230] The data comparison device provided in this application embodiment can be a chip, which includes a processing unit and a communication unit. The processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, pins, or circuits. The processing unit can execute computer execution instructions stored in the storage unit to cause the chip in the server to perform the above-mentioned operations. Figures 4 to 17 The method described in the illustrated embodiment. Optionally, the storage unit is a storage unit within the chip, such as a register, cache, etc. The storage unit can also be a storage unit located outside the chip within the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, random access memory (RAM), etc.

[0231] Specifically, the aforementioned processing unit or processor can be a central processing unit (CPU), a neural-network processing unit (NPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0232] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0233] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0234] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0235] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0236] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. The term "and / or" in this application is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Additionally, the character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not necessarily limited to those steps or modules explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices. The naming or numbering of steps in this application does not imply that the steps in the method flow must be executed in the time / logical order indicated by the naming or numbering. The execution order of the named or numbered process steps can be changed according to the technical purpose to be achieved, as long as the same or similar technical effect can be achieved. The division of modules in this application is a logical division. In actual applications, there may be other division methods. For example, multiple modules may be combined into or integrated into another system, or some features may be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the modules shown or discussed may be through some ports, and the indirect coupling or communication connection between modules may be electrical or other similar forms, which are not limited in this application. Furthermore, the modules or sub-modules described as separate components may or may not be physically separated, may or may not be physical modules, or may be distributed in multiple circuit modules. Some or all of the modules can be selected to achieve the purpose of the solution in this application according to actual needs.

Claims

1. An image processing method, characterized in that, include: Acquire an input image, which is an image captured by an image sensor; Acquire a hyperspectral image, wherein the spectral channels of the hyperspectral image are not exactly the same as those of the input image; Features are extracted from the input image to obtain the first feature; The second feature is obtained based on the features extracted from the hyperspectral image; By fusing the first feature and the second feature, a fused feature is obtained; The output image is obtained based on the fusion features.

2. The method according to claim 1, characterized in that, The fusion of the first feature and the second feature to obtain the fused feature includes: The first feature and the second feature are fused based on an attention mechanism to obtain the fused feature.

3. The method according to claim 1 or 2, characterized in that, The process of fusing the first feature and the second feature to obtain the fused feature further includes: Acquire historical frame images, where the historical frame images are the previous frame images of the input images; Features are extracted from the historical frames to obtain the third feature; The first feature, the second feature, and the third feature are fused together to obtain the fused feature.

4. The method according to any one of claims 1-3, characterized in that, The step of obtaining the second feature based on the features extracted from the hyperspectral image includes: Information about spectral channels is filtered from the hyperspectral image to obtain a reconstructed image of the spectral channels; The second feature is obtained by reconstructing the image based on the features extracted from the spectral channels.

5. The method according to claim 4, characterized in that, The step of filtering information from the hyperspectral image for a preset spectral channel to obtain a spectral channel reconstructed image includes: Features are extracted from the hyperspectral image to obtain the third feature; Based on the attention mechanism, dynamic attention coefficients and spectral features are obtained according to the third feature; The spectral channel reconstructed image is obtained by fusing the dynamic attention coefficient, the spectral features, and the preset static attention coefficient.

6. The method according to claim 4 or 5, characterized in that, The step of obtaining the second feature based on the features extracted from the reconstructed image of the spectral channels includes: The reconstructed spectral channel image is downsampled at least once to obtain at least one downsampled spectral image; Features are extracted from the at least one frame of downsampled spectral image to obtain at least one scale feature, wherein the second feature includes the at least one scale feature.

7. The method according to claim 6, characterized in that, The process of fusing the first feature and the second feature to obtain the fused feature further includes: The first feature is downsampled at least once to obtain at least one downsampled feature, wherein the scale of the at least one downsampled feature corresponds one-to-one with the scale of the at least one scaled feature; The at least one downsampled feature is fused with the corresponding scale feature to obtain at least one scale fused feature; The at least one scale fusion feature is fused to obtain the fused feature.

8. The method according to any one of claims 1-7, characterized in that, The method further includes: Aberration estimation is performed on the hyperspectral image to obtain an aberration estimation map, which is used to represent the error of pixels in the input image. The step of obtaining the output image based on the fusion features includes: The fusion features are recovered based on the aberration estimation map to obtain the output image.

9. The method according to claim 8, characterized in that, The step of performing aberration estimation on the input image based on the hyperspectral image to obtain an aberration estimation map includes: The hyperspectral image is used as input to a pre-trained generator to output the aberration estimation map. Alternatively, an offline calibration algorithm can be used to estimate the aberrations of the input image based on the hyperspectral image to obtain an aberration estimation map. The offline calibration algorithm is an algorithm determined in an offline state according to a pre-set set of steps.

10. The method according to any one of claims 1-9, characterized in that, The process of fusing the first feature and the second feature to obtain the fused feature further includes: The first feature and the second feature are aligned, and the aligned first feature and the second feature are fused to obtain the fused feature.

11. An image processing apparatus, characterized in that, include: The first input module is used to acquire an input image, which is an image captured by an image sensor. The second input module is used to acquire a hyperspectral image, wherein the spectral channels of the hyperspectral image are not exactly the same as the spectral channels of the input image; The first feature extraction module is used to extract features from the input image to obtain a first feature; The second feature extraction module is used to obtain a second feature based on the features extracted from the hyperspectral image; The feature fusion module is used to fuse the first feature and the second feature to obtain a fused feature; The recovery module is used to obtain the output image based on the fusion features.

12. The apparatus according to claim 11, characterized in that, The feature fusion module is specifically used to fuse the first feature and the second feature based on an attention mechanism to obtain the fused feature.

13. The apparatus according to claim 11 or 12, characterized in that, The feature fusion module is also used for: Acquire historical frame images, where the historical frame images are the previous frame images of the input images; Features are extracted from the historical frames to obtain the third feature; The first feature, the second feature, and the third feature are fused together to obtain the fused feature.

14. The apparatus according to any one of claims 11-13, characterized in that, The second feature extraction module is specifically used for: Information about spectral channels is filtered from the hyperspectral image to obtain a reconstructed image of the spectral channels; The second feature is obtained by reconstructing the image based on the features extracted from the spectral channels.

15. The apparatus according to claim 14, characterized in that, The second feature extraction module is specifically used for: Features are extracted from the hyperspectral image to obtain the third feature; Based on the attention mechanism, dynamic attention coefficients and spectral features are obtained according to the third feature; The spectral channel reconstructed image is obtained by fusing the dynamic attention coefficient, the spectral features, and the preset static attention coefficient.

16. The apparatus according to claim 14 or 15, characterized in that, The second feature extraction module is specifically used for: The reconstructed spectral channel image is downsampled at least once to obtain at least one downsampled spectral image; Features are extracted from the at least one frame of downsampled spectral image to obtain at least one scale feature, wherein the second feature includes the at least one scale feature.

17. The apparatus according to claim 16, characterized in that, The feature fusion module is also used for: The first feature is downsampled at least once to obtain at least one downsampled feature, wherein the scale of the at least one downsampled feature corresponds one-to-one with the scale of the at least one scaled feature; The at least one downsampled feature is fused with the corresponding scale feature to obtain at least one scale fused feature; The at least one scale fusion feature is fused to obtain the fused feature.

18. The apparatus according to any one of claims 11-17, characterized in that, The device further includes: An aberration estimation module is used to perform aberration estimation based on the hyperspectral image to obtain an aberration estimation map, which is used to represent the error of pixels in the input image. The recovery module is specifically used to recover the fusion features based on the aberration estimation map to obtain the output image.

19. The apparatus according to claim 18, characterized in that, The aberration estimation module is specifically used for: The hyperspectral image is used as input to a pre-trained generator to output the aberration estimation map. Alternatively, an offline calibration algorithm can be used to estimate the aberrations of the input image based on the hyperspectral image to obtain an aberration estimation map. The offline calibration algorithm is an algorithm determined in an offline state according to a pre-set set of steps.

20. The apparatus according to any one of claims 11-19, characterized in that, The feature fusion module is also used for: The first feature and the second feature are aligned, and the aligned first feature and the second feature are fused to obtain the fused feature.

21. An image processing apparatus, characterized in that, The method includes one or more processors coupled to a memory storing a program that, when executed by the one or more processors, implements the steps of any one of claims 1 to 10.

22. A computer-readable storage medium, characterized in that, The program, when executed by the processing unit, performs the method as described in any one of claims 1 to 10.

23. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method as described in any one of claims 1 to 10.