Hyperspectral image super-resolution method and device based on physical diffusion model

By introducing a physical diffusion model and adaptive physical information guidance into the hyperspectral image super-resolution method, and combining edge and semantic information, the problem of the underutilization of the physical properties of spectral images in existing methods is solved, achieving high-quality high-resolution reconstruction and improving spatial details and spectral fidelity.

CN120634865BActive Publication Date: 2025-10-24HANGZHOU DIANZI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511130271.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-10-24
Estimated Expiration
2045-08-13

AI Technical Summary

Technical Problem

Existing hyperspectral image super-resolution methods do not fully consider the physical characteristics of spectral images, resulting in insufficient reconstruction results in complex scenes and affecting their application in high-precision quantitative analysis.

Method used

A method based on the physical diffusion model is adopted. By introducing adaptive physical information guidance, combining edge information and semantic information, an iterative process of the diffusion model is constructed, and the physical imaging model is used for iterative optimization to reconstruct high-resolution hyperspectral images.

Benefits of technology

It significantly improves the quality of high-resolution hyperspectral image generation, enhances spatial detail clarity and spectral fidelity, and strengthens the physical consistency and robustness of reconstruction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634865B_ABST
    Figure CN120634865B_ABST
Patent Text Reader

Abstract

The application discloses a hyperspectral image super-resolution method and device based on a physical diffusion model. The application acquires a low-resolution hyperspectral image and a corresponding high-resolution multispectral image; extracts edge information and semantic information from the high-resolution multispectral image; inputs the low-resolution hyperspectral image, the edge information and the semantic information into a physical diffusion model to obtain a high-resolution hyperspectral image. The application successfully integrates the physical mechanism and scene physical properties of hyperspectral imaging into the core iteration process of the diffusion model through physical information guiding subnetwork design, physical information extraction and fusion methods and physical model guiding subnetwork design. The integration significantly improves the spatial detail definition, spectral fidelity and overall physical consistency of the high-resolution hyperspectral image, and provides more effective physical guiding information and a more robust physical constraint optimization mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of computer vision and image reconstruction, and particularly relates to a hyperspectral image super-resolution method and device based on a physical diffusion model. BACKGROUND

[0002] Hyperspectral images capture the spectral information of an object with hundreds of continuous bands, providing unique spectral features for material identification, and are applied in precision agriculture, environmental monitoring and other fields. However, due to the physical defects of imaging sensors, it is difficult for hyperspectral imaging systems to maintain high spectral and spatial resolution at the same time. Such low spatial resolution will severely limit the performance in small crop classification and mineral microstructure analysis scenarios. In order to solve the problems caused by this physical limitation, hyperspectral super-resolution technology has become a key.

[0003] Traditional hyperspectral image super-resolution technology is mainly built on the following mathematical theoretical framework: in the field of statistical learning, the Bayesian probability model is based on the statistical correlation between observed data and target images, and the image reconstruction is realized through an iterative optimization algorithm; non-negative matrix factorization decouples spectral features and spatial information, and adopts a hierarchical optimization strategy for bimodal reconstruction; the CP decomposition technology based on tensor algebra fully utilizes the three-dimensional characteristics of hyperspectral data to realize spatial-spectral joint feature extraction; the method based on sparse representation constructs a dual-resolution dictionary mapping system, and realizes resolution enhancement through cross-scale feature matching. However, traditional methods rely too much on idealized assumptions and manually designed features, which limits the robustness, efficiency and generalization ability in complex real-world scenarios.

[0004] The development of deep learning methods has promoted progress in various fields, and hyperspectral image super-resolution has also benefited greatly. Neural networks, especially Convolutional Neural Network (CNN), Transformer and Generative Adversarial Network (GAN), have been relatively mature in the application of hyperspectral image super-resolution, and these models still face common and difficult challenges. CNN shows excellent reconstruction performance, but its internal structure often lacks transparency and theoretical support. Although Transformer is good at processing long sequence data, the model complexity is high and a large amount of training data is needed. GAN models have proven their ability to generate high-quality results, but they are plagued by problems such as training instability, lack of interpretability and pattern collapse susceptibility, which greatly affect the quality of super-resolution hyperspectral images.

[0005] In recent years, denoising diffusion probability model (DDPM) is known for its high-quality generation, strong interpretability and high flexibility. Therefore, using DDPM in spectral image super-resolution can be a solution to alleviate the above model dilemma. For example, Li et al. introduced DCDM (Dual Conditional Diffusion Model), which includes two noise prediction networks with different conditional inputs, specifically for extracting image features from low-resolution spectral images and panchromatic images, adapting to different features of input images, and reconstructing high-resolution spectral image features for fusion. However, these methods directly combine DDPM and spectral image super-resolution, ignoring the constraints and guiding effects of the physical characteristics of spectral images on the reconstruction process, which may lead to deficiencies in the reconstruction results and limit their application effect in scenarios requiring high-precision quantitative analysis. SUMMARY

[0006] The present application aims to solve the problem that existing hyperspectral image super-resolution methods based on diffusion models do not fully consider the physical characteristics of spectral images. A hyperspectral image super-resolution method and device based on a physical diffusion model are proposed, which introduces adaptive physical information guidance in the iteration of the diffusion model to improve the reconstruction quality.

[0007] To achieve the above purpose, the present application adopts the following technical solutions.

[0008] In a first aspect, the present application provides a hyperspectral image super-resolution method based on a physical diffusion model, which comprises:

[0009] obtaining a low-resolution hyperspectral image and a corresponding high-resolution multispectral image ; extracting edge information and semantic information from the high-resolution multispectral image ;

[0010] inputting the low-resolution hyperspectral image , edge information , and semantic information to a physical diffusion model to obtain a high-resolution hyperspectral image .

[0011] Preferably, the physical diffusion model includes a forward Markov chain process and a reverse Markov chain process.

[0012] The reverse Markov chain process includes multiple stages; wherein the first stage gradient term merges observation constraints into the generated image at the current time step through a noise prediction network . The noise prediction network comprises a physical information guided spectral denoising subnetwork and a physical model guided spectral denoising subnetwork in series.

[0013] As preferably, the training loss function of the physical diffusion model is represented as:

[0014] Formula (1)

[0015] wherein represents a high-resolution hyperspectral image label, represents the last stage output predicted by the reverse Markov chain process under the stage .

[0016] As preferably, the physical information guided spectral denoising subnetwork is implemented as:

[0017] Formula (2)

[0018] wherein represents the concatenation of multiple information in the channel dimension, represents a mapping function of the physical information guided spectral denoising subnetwork, represents a low-resolution hyperspectral image, represents edge information, represents semantic information, represents a noisy high-resolution hyperspectral image of the stage .

[0019] As preferably, the physical information guided spectral denoising subnetwork uses a Unet network.

[0020] As preferably, the physical model guided spectral denoising subnetwork comprises K stages; wherein the implementation process of each stage is:

[0021] Formula (3)

[0022] Formula (4)

[0023] wherein and are learnable parameters, is a neural network, ∈[0,K-1], B represents a spatial downsampling operation, and S represents a spectral downsampling operation, represents the output of the k+1th stage, adopts a noisy high-resolution hyperspectral image of the stage . ​

[0024] As preferred, the neural network in the physical model guided spectral denoising sub-network is composed of a first convolutional layer, an activation layer, and a second convolutional layer.

[0025] In a second aspect, the present application provides a hyperspectral image super-resolution device, comprising:

[0026] a data acquisition module, responsible for acquiring a low-resolution hyperspectral image and a corresponding high-resolution multispectral image ; extracting edge information and semantic information from the high-resolution multispectral image ;

[0027] an image reconstruction module, responsible for inputting the low-resolution hyperspectral image , edge information , and semantic information to a physical diffusion model to reconstruct a high-resolution hyperspectral image .

[0028] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed in a computer, the computer executes the method.

[0029] In a fourth aspect, the present application provides a computing device comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method is implemented.

[0030] The beneficial effects of the present application at least include:

[0031] (1) This paper proposes a hyperspectral image super-resolution method based on a physical diffusion model. This method constructs a diffusion model through a physical information guidance subnetwork and a physical model guidance subnetwork, thereby improving the generation quality of high-resolution hyperspectral images while making full use of physical information. Specifically, the physical information guidance subnetwork explicitly uses the edge information and semantic information extracted from the high-resolution multispectral image as conditional input, together with the low-resolution hyperspectral image and the current noisy state. This enables the denoising process to adaptively utilize the inherent spatial structure and semantic category information of the scene, guiding the model to generate high-quality spatial details that are more consistent with the spatial structure and semantic constraints of the real scene. The physical model guidance subnetwork solves the physical imaging model by unfolding the optimization and fusing the learnable parameters and neural network. In each iteration, the reconstruction result is forcibly constrained to a feasible solution space that conforms to the imaging physical process (defined by spatial downsampling B and spectral downsampling S). This dual guidance mechanism (physical property guidance + physical model constraints) is deeply coupled to the inverse denoising chain of the diffusion model, ensuring that the final reconstructed high-resolution hyperspectral image not only has high visual quality, but also is consistent with the edge structure of the input high-resolution multispectral image in spatial details, and satisfies the constraints of the degradation model in the spectral dimension, significantly improving the physical consistency and spectral fidelity of the reconstruction results.

[0032] (2) The present invention proposes a physical information extraction method that combines edge detection and semantic segmentation. This method can extract physical information from high-resolution multispectral images to improve the quality of high-resolution hyperspectral image generation of the denoising diffusion model. Specifically, edge information provides a strong prior for high-frequency spatial details such as object contours and texture boundaries; the extracted semantic information provides pixel-level labels for different category areas in the scene. By inputting these high-confidence physical information as conditions into the physical information guidance subnetwork, the diffusion model is provided with richer and more accurate scene spatial structure and semantic context guidance than a single low-resolution image Y in each step of the denoising process. This method effectively overcomes the problems of spatial structure ambiguity or unreasonable semantics that may arise from relying solely on data-driven learning, enabling the model to generate high-resolution spectral images with clearer spatial details and more reasonable semantics.

[0033] (3) This paper proposes an iterative optimization mechanism that integrates a physical imaging model to improve reconstruction accuracy and algorithm robustness. The physical model guides the spectral denoising subnetwork to explicitly construct the physical degradation model of hyperspectral imaging as an objective function and embeds it into each denoising step of the diffusion model through a K-stage iterative optimization process. This approach of embedding the physical model as an iterative optimization module into the diffusion model continuously imposes hard constraints during the generation process, significantly improving the accuracy of the reconstruction results to the imaging physical process and enhancing the adaptability to different degradation conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 This is a general flow chart of the hyperspectral image super-resolution method based on the physical diffusion model disclosed in the present invention.

[0035] Figure 2 This is a schematic diagram of the principle of spectral denoising based on physical information in the present invention.

[0036] Figure 3 This is a schematic diagram of the principle of spectral denoising based on the physical model in the present invention.

[0037] Figure 4 This is a schematic diagram of the principle of iteratively reconstructing a high-resolution hyperspectral image using a diffusion model in the present invention.

[0038] Figure 5 These are diagrams showing the application effects of the present invention on different datasets, where (a) is the ICVL dataset, (b) is the Harvard dataset, and (c) is the Chikusei dataset. DETAILED DESCRIPTION

[0039] This paper discloses a hyperspectral image super-resolution method based on a physically guided diffusion model. It can uniformly process low-resolution hyperspectral images from different datasets and generate diverse high-resolution hyperspectral images, expanding the application scope of hyperspectral image super-resolution. The method has applications in various fields, including precision agriculture and environmental monitoring.

[0040] In order to better illustrate the advantages of the present invention, the invention is further described below with reference to the accompanying drawings and examples.

[0041] like Figure 1 As shown, this embodiment discloses a hyperspectral image super-resolution method based on a physical diffusion model, including:

[0042] Step S1: Generate a low-resolution hyperspectral image for testing using a real high-resolution hyperspectral image and the corresponding high-resolution multispectral image .

[0043] In one implementation, this example uses hyperspectral images from the publicly available ICVL, Harvard, and Chikusei datasets. Each hyperspectral image retains the number of channels, and is cropped and normalized in the spatial dimension to produce a 256×256 spatial resolution hyperspectral image. The low-resolution hyperspectral image is obtained by applying a 9×9 Lanczos blur to the cropped image and downsampling it by a factor of 8. The high-resolution multispectral image is obtained by multiplying the cropped image by the spectral response function.

[0044] Step S2: edge detection operator from high-resolution multispectral image extracting edge information ;

[0045] In one embodiment, the edge detection operator uses a Sobel operator. High-resolution hyperspectral images and high-resolution multispectral images in the same scene, although different in spectral dimension, share consistent spatial edge structures (e.g., object contours, texture boundaries). As shown in FIG. 1, it can be specifically expressed as: Figure 2

[0046] Equation (1)

[0047] Equation (2)

[0048] Equation (3)

[0049] wherein represents a longitudinal gradient operator, represents a transverse gradient operator, represents a gradient intensity of the high-resolution multispectral image, i.e., an edge intensity.

[0050] Step S3: extracting semantic information from the high-resolution multispectral image by a semantic segmentation operator ;

[0051] In one embodiment, the semantic segmentation operator uses SAM (Segment Anything Model), which can obtain an abnormally detailed semantic segmentation mask of a given image, and the process is expressed as:

[0052]

[0053] wherein, represents a semantic segmentation operator, represents the extracted segmentation semantic information;

[0054] Step S4: inputting the low-resolution hyperspectral image , the edge information , and the semantic information to a physical diffusion model to reconstruct a high-resolution hyperspectral image .

[0055] The physical diffusion model includes a forward Markov chain process and a reverse Markov chain process;

[0056] The first stage of the forward Markov chain process is expressed as:

[0057] Equation (4)​​​

[0058] where is the given hyper-parameter, , denotes the noisy high-resolution hyperspectral image at the th stage, denotes the noise term following the standard normal distribution; denotes the high-resolution hyperspectral image label in the training dataset, denotes the noise term at the t-1th step.

[0059] The reverse Markov chain process includes multiple stages;

[0060] The th stage gradient term represents the incorporation of observation constraints into the process of generating images at the current time step as:

[0061] Equation (5)

[0062] where represents a pre-trained noise prediction network;

[0063] The last stage output predicted by the reverse Markov chain process at the th stage is represented as:

[0064] Equation (6)

[0065] The training loss function is represented as:

[0066] Equation (7)

[0067] As Figure 4 described, the noise prediction network includes a physical information guided spectral denoising subnetwork and a physical model guided spectral denoising subnetwork in series;

[0068] As Figure 2 described, the implementation process of the physical information guided spectral denoising subnetwork is as follows:

[0069] Equation (8)

[0070] where denotes the concatenation of multiple information in the channel dimension, denotes the mapping function of the physical information guided spectral denoising subnetwork, is the high-resolution multispectral image.

[0071] As an example, the physical information guided spectral denoising subnetwork can use a Unet network. ​

[0072] The physical model guided spectral denoising subnetwork is developed from the HQS optimization algorithm to make full use of high-resolution multispectral image and low-resolution hyperspectral image information.

[0073] Formula (9)

[0074] The target function is solved by HQS and can be divided into two steps: And Wherein The first step can be solved by gradient descent method, And the second step can be solved by neural network.

[0075] As Figure 3 The physical model guided spectral denoising subnetwork comprises K stages; wherein the implementation process of each stage is:

[0076] Formula (10)

[0077] Formula (11)

[0078] Wherein And Are learnable parameters, Is a neural network, ∈[0,K-1], B represents a spatial down-sampling operation, S represents a spectral down-sampling operation, Indicates the output of the k+1 stage. = As the output of the physical model guided spectral denoising subnetwork. As an example,

[0079] The neural network is composed of a first convolutional layer, an activation layer and a second convolutional layer. In the F update step, the gradient is directly calculated by using the observed data (Z, Y) and the current estimation (F, V)

[0080] , The intermediate solution is forced to be adjusted in the direction conforming to the physical observation. In the V update step, a light neural network D is used to impose image prior. The introduction of learnable parameters (ρ, α) enables the model to adaptively adjust the weights of the physical constraint term and the prior term in the optimization process, and adapt to different degradation intensities and image contents. The hyperspectral image super-resolution method based on the physical diffusion model can process different low-resolution hyperspectral images and reconstruct high-resolution hyperspectral images.

[0081] (a) to Figure 5 (a) to Figure 5The middle (c) shows the super-resolution effect of the present application on the ICVL dataset, the Harvard dataset and the Chikusei dataset. The hyperspectral image super-resolution method of the present application can reconstruct high-quality high-resolution hyperspectral images.

[0082] To further illustrate the effect of the present application, the present embodiment compares various methods under the same experimental conditions. The experimental results are shown in the following table:

[0083]

[0084] As can be seen from the results in the table, the method disclosed in the present application can achieve very good super-resolution effect, and the PSNR, SSIM and ERGAS indicators are significantly higher than those of the comparison methods on multiple datasets. PSNR and SSIM mainly measure the spatial quality of the hyperspectral image after super-resolution, while ERGAS mainly measures the spectral quality of the hyperspectral image after super-resolution. Therefore, the hyperspectral image obtained by the method disclosed in the present application has smaller spatial error and spectral error, and is superior to other methods in terms of spatial quality and spectral fidelity.

[0085] The present embodiment also discloses a hyperspectral image super-resolution device, comprising:

[0086] a data acquisition module responsible for acquiring low-resolution hyperspectral images and corresponding high-resolution multispectral images ; extracting edge information and semantic information from the high-resolution multispectral images ;

[0087] an image reconstruction module responsible for inputting the low-resolution hyperspectral images , the edge information , the semantic information to a physical diffusion model to reconstruct high-resolution hyperspectral images .

[0088] The present embodiment also discloses a computing device, specifically, the computing device comprises a memory and a processor, the memory stores executable code, and the processor executes the executable code to implement the method of any one of the embodiments.

[0089] The storage component can be configured with a high-speed random access memory (RAM) and can also be extended with a non-volatile storage unit (such as at least one disk storage device). Through at least one transmission interface (supporting wired / wireless communication mode), a data channel is established between the device node and other nodes, and a multi-level network architecture (including the Internet, a wide area network, a local area network and a metropolitan area network) is adapted.

[0090] The bus system can be of the type of industry standard architecture (ISA) bus, peripheral component interconnect (PCI) bus or extended industry standard architecture (EISA) bus, etc., and its internal architecture is divided into three types of functional modules, namely address transmission channel, data transmission channel and control signal channel.

[0091] The storage component is used for storing executable codes, and when the operation unit receives an execution instruction, the codes are executed to implement the method flow of the embodiments of the application.

[0092] The operation unit can adopt a semiconductor integrated circuit architecture and has digital signal analysis and operation capability. The method implementation includes two paths: direct hardware decoding through a physical layer transistor logic circuit or soft / hardware cooperative operation based on a programmable instruction set. In a specific implementation, the operation unit can be adapted to multiple types of computing architectures: a general-purpose computing unit (including CPU / NP suitable for scalar instruction processing; a heterogeneous computing unit (including DSP / ASIC / FPGA) supporting customized algorithm acceleration; and a reconfigurable computing unit (including discrete gate circuits / transistor arrays) providing hardware-level dynamic configuration capability. The instruction execution carrier is implemented through a storage hierarchy system: a software instruction set is stored in a non-volatile storage matrix (including NOR Flash / EEPROM) or a high-speed temporary storage area (SRAM / register stack) of the storage component, and is accessed and scheduled by the operation unit through a bus architecture. The hardware acceleration module can be directly mapped to an on-chip storage medium (Cache / BRAM) to realize zero-delay response of the instruction.

[0093] The embodiment of the application provides a readable storage medium in the form of a computer program product, which contains a computer readable storage carrier with a program code storage function, the program code loaded in the carrier contains executable instructions for implementing the technical solutions in the foregoing method embodiments, and the specific implementation details of the related method steps have been completely recorded in the foregoing method embodiments, and the embodiment does not repeat the description.

[0094] When the functional module is implemented in the form of an independent software unit and commercialized as a commodity, its running data can be stored in a digital information storage medium. Based on the technical implementation principle, the innovative core value or the feature module that is different from the prior art of the patent can be presented in the form of an application program package. Such a program suite is usually stored in a data carrier and has multiple operation commands built-in to guide an electronic device (including personal terminals, cloud servers, networked devices and other hardware facilities) to perform all or core operation links of the detailed processes in the patent embodiments. The information carrier includes a portable flash disk, an external storage device, a firmware memory (ROM), a dynamic access memory (RAM), a magnetic storage disk or an optical recording medium, etc.

[0095] The above detailed description of the specific embodiments of the present application, the purpose, technical solutions and beneficial effects of the present application are further described in detail, and it should be understood that the above description is only a specific embodiment of the present application and does not limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A hyperspectral image super-resolution method based on a physical diffusion model, characterized in that, The method comprises: Acquiring low-resolution hyperspectral images and the corresponding high-resolution multispectral image ; From high-resolution multispectral images Extract edge information and semantic information ; A low resolution hyperspectral image , edge information , semantic information is input to a physical diffusion model, and a high resolution hyperspectral image is reconstructed ; The physical diffusion model comprises a forward Markov chain process and a reverse Markov chain process; The reverse Markov chain process includes multiple stages; wherein, the first stage The stage gradient term merges the observation constraint into the generated image at the current time step by a noisy prediction network; the noisy prediction network includes a physical information guided spectral denoising sub-network and a physical model guided spectral denoising sub-network in series. The physical model guided spectral denoising sub-network comprises K stages; wherein the implementation process of each stage is: Equation (3) Equation (4) wherein and are learnable parameters, is a neural network, ∈ [0, K-1], B denotes a spatial downsampling operation, S denotes a spectral downsampling operation, denotes the output of the k+1th stage, the first stage of the noisy high-resolution hyperspectral image .

2. The method of claim 1, wherein, The training loss function of the physical diffusion model is represented as: Formula (1) wherein represents a high-resolution hyperspectral image label, represents the last stage output predicted by the reverse Markov chain process under the stage.

3. The method of claim 1, wherein, The physical information guides the spectrum denoising sub-network The implementation process is: Formula (2) wherein represents concatenation of multiple information in channel dimension, represents a mapping function of the physically information guided spectral denoising subnetwork, represents a low resolution hyperspectral image, represents edge information, represents semantic information, represents a first stage of a noisy high resolution hyperspectral image.

4. The method of claim 3, wherein, The physical information guided spectral denoising sub-network uses an Unet network.

5. The method of claim 1, wherein, The neural network in the physical model guided spectral denoising sub-network is composed of a first convolutional layer, an activation layer, and a second convolutional layer.

6. A hyperspectral image super-resolution device for implementing the method of any one of claims 1-5, characterized in that, Comprise: a data acquisition module, responsible for acquiring low-resolution hyperspectral images and corresponding high-resolution multispectral images ; extracting edge information and semantic information from high-resolution multispectral images ​ An image reconstruction module is responsible for inputting the low-resolution hyperspectral image , edge information , semantic information to a physical diffusion model to reconstruct a high-resolution hyperspectral image .

7. A computer readable storage medium having stored thereon a computer program, which, when executed in a computer, causes the computer to carry out the method of any one of claims 1-5.

8. A computing device comprising a memory and a processor, the memory having stored therein executable code, the processor executing the executable code to implement the method of any one of claims 1-5.

Citation Information

Patent Citations

  • Hyperspectral and multispectral image fusion method and system based on self-learning coupling diffusion posterior sampling

    CN118628366A