Chained diffusion remote sensing hyperspectral image super-resolution system and method

The chain diffusion system addresses the challenges of high-spectral resolution imaging by integrating dynamic graph learning and diffusion models to enhance spatial and spectral fidelity, achieving improved reconstruction quality and robustness in complex regions.

CN120318077AActive Publication Date: 2025-07-15CHANGCHUN INST OF OPTICS FINE MECHANICS & PHYSICS CHINESE ACAD OF SCI

Patent Information

Application Number
CN202510799513.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-07-15
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

Current methods for high-spectral resolution imaging face challenges in capturing high-definition spatial details due to hardware limitations, and existing models struggle with complex spectral-spatial coupling, leading to issues like mixed pixel effects and texture fuzziness in high-spectral resolution images.

Method used

A chain diffusion-based system combining dynamic graph learning with diffusion models, utilizing time-varying weights for iterative optimization, and a differential-frequency collaborative attention mechanism to balance global and local modeling, enhanced by a semantic constraint loss function for improved image reconstruction.

Benefits of technology

The system achieves enhanced spatial resolution and spectral fidelity by jointly modeling noise processes and feature evolution, improving reconstruction quality in complex regions and maintaining global trends with local patterns, while being robust to noise and distortions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318077A_ABST
    Figure CN120318077A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of remote sensing image processing, in particular to a chain diffusion remote sensing hyperspectral image super-resolution system and method, in a dynamic hypergraph diffusion branch, dynamic hypergraph learning and a diffusion model are combined, an iterative optimization process is guided through a time-varying weight space, and joint modeling of a denoising process and feature evolution is realized; in the difference-frequency collaborative attention branch, a difference-frequency collaborative attention mechanism is constructed, the limitation of traditional spectrum-space separation modeling is broken through, and balance is achieved between global trend modeling and local form keeping; according to the method provided by the invention, by constructing the semantic constraint loss function and through a semantic-driven local constraint mechanism, the reconstruction quality of the complex region is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing image processing, and particularly relates to a hyperspectral remote sensing image super-resolution system and method with chain diffusion. Background Art

[0002] With its unique multi-band spectral information capture characteristics, hyperspectral remote sensing technology can provide accurate spectral feature support for ground object recognition and classification. Limited by the hardware conditions of existing imaging systems, it is often difficult to obtain high-definition spatial detail information during the acquisition of hyperspectral images. This lack of spatial resolution severely restricts the practical application value of this technology in application requirements such as fine ground object recognition and weak target detection.

[0003] The core scientific problem of hyperspectral image super-resolution lies in achieving the dual optimization goals of spatial detail reconstruction and spectral feature preservation. With the breakthrough progress of machine learning technology, deep learning has shown superior performance in the field of computer vision, especially forming a new technical paradigm in the field of image quality enhancement. Compared with traditional physical model-based optimization methods, deep neural networks achieve better spatial detail reconstruction while maintaining spectral fidelity by autonomously mining high-order non-linear features in the data. Different from traditional RGB images, the requirement for spectral fidelity in reconstructing hyperspectral images greatly increases the challenge of this task. 2D convolution has good performance in capturing spatial information, while 3D convolution has unique advantages in spectral feature mining. However, this special three-dimensional operation will cause an explosive growth in computational complexity, thereby restricting the application scenarios of the model. Thanks to its outstanding global modeling ability, Transformer and its variants have shown breakthrough performance in image processing tasks. Compared with the CNN (Convolutional Neural Network) framework that is good at extracting local information, the global self-attention-based Transformer can establish long-range dependencies between global pixels within a single layer network. Remote sensing images usually correspond to a large observation area. Therefore, the modeling ability for global information becomes a key factor in improving the reconstruction quality. In the field of hyperspectral image super-resolution, a common approach is to combine Transformer with a 3D convolutional network to construct a learning framework to achieve collaborative modeling of global and local information. However, the Transformer-based framework lacks the guidance of explicit prior information during the learning process, which greatly limits the learning efficiency of the model.

[0004] The current research on hyperspectral remote sensing image super-resolution faces the following challenges: 1) The complex spectral-spatial coupling relationship presented in remote sensing images poses higher requirements for the high-order modeling ability of the model. Remote sensing images usually correspond to a large observation area and a low spatial resolution, which results in a large number of mixed pixels in the image. There is a strong correlation between the spectral reflection characteristics and spatial distribution of mixed pixels. Traditional convolutional neural networks are difficult to capture global dependencies, while Transformer lacks the guidance of explicit prior information; 2) The lack of joint constraints on multi-level features often leads to blurred textures in complex regions. The effective guidance of spectral and spatial structure characteristics at multiple scales for the learning process is of great significance for maintaining semantic and structural consistency. The lack of joint constraints often leads to spectral distortion and local texture blurring. Summary of the Invention

[0005] In view of this, the present invention aims to provide a chain diffusion-based remote sensing hyperspectral image super-resolution system and method, which combines dynamic hypergraph learning with a diffusion model, and realizes the joint modeling of the denoising process and feature evolution by guiding the iterative optimization process with a time-varying weight space; through a differential-frequency collaborative attention mechanism, it breaks through the limitations of traditional spectral-spatial separated modeling and achieves a balance between global trend modeling and local morphology preservation; the proposed semantic constraint loss function significantly improves the reconstruction quality of complex regions through a semantic-driven local constraint mechanism.

[0006] To achieve the above object, the technical solution of the present invention is realized as follows: A chain diffusion-based remote sensing hyperspectral image super-resolution system, comprising: A dynamic hypergraph diffusion branch, which embeds the time condition of the input image into the input image, updates and learns the dynamic adjacency matrix of the obtained features, and obtains affine features based on the dynamic adjacency matrix; denoises and extracts details from the affine features to obtain diffusion features; A differential-frequency collaborative attention branch, which performs a two-branch spatial attention operation on the input image, completes differential and frequency analysis, and respectively obtains differential features and frequency features; combines the differential features and frequency features to obtain collaborative features; A super-resolution reconstruction branch, which combines the diffusion features, collaborative features, and the upsampled image of the input image to obtain a super-resolution image; Feature fusion is performed through mutual attention operations between the dynamic hypergraph diffusion branch and the differential-frequency collaborative attention branch, and the fused features are used to respectively guide the dynamic hypergraph diffusion branch and the differential-frequency collaborative attention branch for feature processing.

[0007] Furthermore, the dynamic hypergraph diffusion branch includes multiple cascaded groups of dynamic hypergraph diffusion units; each group of dynamic hypergraph diffusion units includes two cascaded dynamic hypergraph diffusion units. The output features of the first dynamic hypergraph diffusion unit are input into the mutual attention operation, and the output features of the mutual attention operation are added element-wise to the output features of the second dynamic hypergraph diffusion unit. In each dynamic hypergraph diffusion unit: According to the input image or input features, obtain the corresponding basic hypergraph adjacency matrix; Add noise to the input image or input features to obtain the diffusion features at the initial time step. At the same time, determine the corresponding time-varying feature vector according to the number of channels of the input image or input features; linearly map the time-varying feature vector combined with the diffusion features at the current time step to obtain the linearly mapped features; perform a query-key attention operation on the linearly mapped features to obtain the similarity matrix; Then calculate the Hadamard product of the similarity matrix and the corresponding basic hypergraph adjacency matrix to obtain the corresponding hypergraph adjacency matrix; Calculate the element-wise product of the hypergraph adjacency matrix, the linearly mapped features, and the learnable projection matrix to obtain the affine features in the current dynamic hypergraph diffusion unit; Denoise and extract features from the affine features, inject the processed features into the input image or input features to obtain the diffusion features at the next time step; repeat the processing of the diffusion features at the current time step for the diffusion features at the next time step to obtain the final diffusion features of the current dynamic hypergraph diffusion unit.

[0008] Furthermore, the process of obtaining the basic hypergraph adjacency matrix in each dynamic hypergraph diffusion unit includes: Through the K-nearest neighbor algorithm, aggregate the K nodes with the highest spectral similarity to each target node in the input image or input features to each hyperedge; Assign weights to each node on the hyperedge through the negative exponential function to obtain the basic hypergraph adjacency matrix.

[0009] Furthermore, the process of determining the corresponding time-varying feature vector according to the number of channels of the input image or input features includes: Through the following formula, use sine time encoding to generate the preliminary time-varying feature vector: ; ; where k represents the serial number of the frequency component, represents the fundamental frequency of the k-th frequency component, d represents the number of channels of the input image or input features, represents the preliminary time-varying feature vector at the t-th iteration; Perform dimensional compression on the initial time-varying feature vector through multiple consecutive MLPs to obtain the time-varying feature vector.

[0010] Further, the process of denoising and feature extraction for the affine features in each dynamic hypergraph diffusion unit includes: performing a convolution operation on the affine features, performing a group normalization operation on the convolved features, and performing a GELU activation operation on the features after the group normalization operation to obtain the output features.

[0011] Further, the process of injecting the processed features into the input image or input features in each dynamic hypergraph diffusion unit includes: ; ; ; where represents the noise attenuation rate, represents the detail injection rate, represents the input image or the current input features, represents the features after denoising and feature extraction processing of the affine features, and LinearSchedule represents a function for linearly adjusting hyperparameters.

[0012] Further, in the differential-frequency collaborative attention branch, after performing convolution feature extraction on the input image to obtain the input features, the input features are processed by multiple cascaded groups of differential-frequency collaborative attention units; where each group of differential-frequency collaborative attention units includes two cascaded differential-frequency collaborative attention units, the output features of the first differential-frequency collaborative attention unit are input into the mutual attention operation, and the output features of the mutual attention operation are added element-wise to the output features of the second differential-frequency collaborative attention unit; each differential-frequency collaborative attention unit includes a differential attention analysis sub-branch and a frequency analysis sub-branch, the input features are processed by the differential attention analysis sub-branch and the frequency analysis sub-branch simultaneously to obtain the differential features and frequency features respectively; the differential features and frequency features are added element-wise, and after the added features are subjected to a convolution operation, they are added element-wise to the input features to obtain the collaborative features; In the differential attention analysis sub-branch, the input features are convolved using multiple groups of spectral differential kernels with different scales to complete the differential calculation of multi-scale spectral dimension sliding; the obtained multi-scale spectral features are processed by a multi-layer perceptron to generate multi-scale attention weights; after the multi-scale attention weights are fused, a convolution fine-tuning operation is performed, and then multiplied element-wise with the input features to obtain the differential features; In the frequency analysis sub-branch, the Laplacian operator and the Gaussian kernel are used to extract high-frequency features and low-frequency features from the input features respectively; after channel concatenation of the high-frequency features and the low-frequency features, a convolution fusion operation is performed on the concatenated features, and then the fused features are multiplied by the corresponding elements of the input features to obtain frequency features.

[0013] Furthermore, in the super-resolution reconstruction sub-branch, the corresponding elements of the diffusion features and the collaborative features are added together, and after continuously performing a convolution operation and a transposed convolution operation on the added features, they are added to the upsampled image element by element; after performing a convolution operation on the added features, a super-resolution image is obtained.

[0014] A super-resolution method for remote sensing hyperspectral images with chain diffusion includes: S1: Obtain a remote sensing image dataset, and perform preprocessing on the remote sensing image dataset to obtain a training set; S2: Input the training set obtained in step S1 into the super-resolution system for remote sensing hyperspectral images with chain diffusion provided by the present invention, and train the super-resolution system in combination with a semantic constraint loss function to obtain a super-resolution model; S3: According to the super-resolution model obtained in step S2, adjust the hyperparameters for training the super-resolution system, and repeat step S2 until an optimal super-resolution model is obtained; S4: Input a low-resolution remote sensing image into the optimal super-resolution model obtained in step S3 for super-resolution reconstruction to obtain a corresponding high-resolution remote sensing image.

[0015] Furthermore, the semantic constraint loss function in step S2 is obtained by the following formula: ; where represents the semantic constraint loss function, represents the L1 loss function, represents the SAM loss function, represents the superpixel nesting loss function, 、 and represent loss weights; The superpixel nesting loss function is obtained by the following formula: ; where represents the spectral loss, represents the texture loss, represents the superpixel nesting loss weight; The spectral loss is obtained by the following formula: ; Among them, represents the number of superpixels obtained after superpixel segmentation of the super-resolution image output by the super-resolution model using the SLIC algorithm, represents the spectral mean of all pixels within the i-th superpixel in the true high-resolution remote sensing image corresponding to the super-resolution image, represents the spectral mean of all pixels within the i-th superpixel in the super-resolution image. The spectral mean is obtained by the following formula: ; Among them, represents the superpixel, represents the number of pixels included in the superpixel, represents the pixel value at the (x, y) position in image I; Texture loss is obtained by the following formula: ; Among them, represents the texture entropy of all pixels within the i-th superpixel in the true high-resolution remote sensing image, represents the texture entropy of all pixels within the i-th superpixel in the super-resolution image. The texture entropy is obtained by the following formula: ; Among them, represents the probability of the b-th bin in the histogram distribution of LBP values within the superpixel, and B represents the number of intervals of the histogram.

[0016] Compared with the prior art, the present invention can achieve the following beneficial effects: (1) In the remote sensing hyperspectral image super-resolution system with chain diffusion of the present invention, by combining dynamic hypergraph learning with a diffusion model and guiding the iterative optimization process through time-varying weight space, the joint modeling of the denoising process and feature evolution is realized; through the differential-frequency collaborative attention mechanism, the limitation of traditional spectral-spatial separation modeling is broken, and a balance is achieved between global trend modeling and local morphology preservation; the proposed semantic constraint loss function significantly improves the reconstruction quality of complex regions through a semantically driven local constraint mechanism; (2) In the chain diffusion-based remote sensing hyperspectral image super-resolution system of the present invention, the asymptotic denoising process of the diffusion model is innovatively combined with the dynamic evolution of the hypergraph structure. Through a time-conditioned hyperedge weight adjustment mechanism, collaborative modeling of the noise statistical characteristics and the spectral-spatial relationship is achieved. The dynamic hypergraph update mechanism allows the model to adaptively adjust the hypergraph node correlation and optimize the learning process according to the characteristic features at different degradation stages. While using noise for information supplementation, the global structural consistency and local detail fidelity are simultaneously optimized, and the anti-disturbance ability of the model is effectively improved; (3) In the chain diffusion-based remote sensing hyperspectral image super-resolution system of the present invention, the multi-scale difference strategy can suppress noise interference and retain high-frequency details by fusing gradient features (such as local mutations and global trends) with different receptive fields, thereby accurately maintaining spectral continuity while improving the spatial resolution; (4) In the chain diffusion-based remote sensing hyperspectral image super-resolution method of the present invention, a superpixel nested loss is proposed. Taking the superpixel as the basic semantic unit, the spectral and texture features are jointly optimized, thereby ensuring that the reconstruction result conforms to the physical semantic laws of the real world while improving the spatial resolution. This design makes up for the deficiencies of traditional pixel-level loss functions in terms of local structure fidelity and semantic consistency, providing a new optimization perspective for high-dimensional image reconstruction tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings: Figure 1 is a schematic diagram of the chain diffusion-based remote sensing hyperspectral image super-resolution system of the present invention according to an embodiment of the present invention; Figure 2 is a schematic diagram of the dynamic hypergraph diffusion unit according to an embodiment of the present invention; Figure 3 is a schematic diagram of the cascaded difference-frequency collaborative attention unit according to an embodiment of the present invention; Figure 4 is a flowchart of the chain diffusion-based remote sensing hyperspectral image super-resolution method of the present invention; Figure 5 is a visualization comparison result diagram of the reconstruction method according to an embodiment of the present invention and other reconstruction methods on the MDAS dataset; Figure 6 is a visualization comparison result diagram of the reconstruction method according to an embodiment of the present invention and other reconstruction methods on the Pavia Centre dataset; Figure 7 This is a visualization comparison result graph of the reconstruction method described in the embodiments of the present invention creation and other reconstruction methods on the Houston dataset. Detailed implementation manners

[0018] In order to make the objectives, technical solutions and advantages of the present invention creation clearer and more understandable, the present invention creation will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention creation and do not constitute a limitation to the present invention creation.

[0019] It should be noted that, without conflict, the embodiments in the present invention creation and the features in the embodiments can be combined with each other.

[0020] In the description of the present invention creation, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention creation and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the present invention creation. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention creation, unless otherwise specified, the meaning of "a plurality" is two or more.

[0021] In the description of the present invention creation, it should be noted that, unless otherwise clearly specified and limited, the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected, or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention creation can be understood through specific situations.

[0022] The present invention creation will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0023] As Figure 1As shown in the figure, the chain diffusion-based remote sensing hyperspectral image super-resolution system of the present invention described in the embodiments of the present invention includes a dynamic hypergraph diffusion branch, a differential-frequency collaborative attention branch, and a super-resolution reconstruction branch. Among them, the dynamic hypergraph diffusion branch embeds the time condition of the input image into the input image, updates and learns the dynamic adjacency matrix of the obtained features, and obtains the affine features based on the dynamic adjacency matrix; denoises and extracts details from the affine features to obtain the diffusion features. The differential-frequency collaborative attention branch performs a two-branch spatial attention operation on the input image, completes differential and frequency analysis, and respectively obtains differential features and frequency features; combines the differential features and frequency features to obtain collaborative features. The super-resolution reconstruction branch combines the diffusion features, collaborative features, and the upsampled image of the input image to obtain the super-resolution image. Feature fusion is performed between the dynamic hypergraph diffusion branch and the differential-frequency collaborative attention branch through mutual attention operation, and the fused features are used to respectively guide the dynamic hypergraph diffusion branch and the differential-frequency collaborative attention branch for feature processing.

[0024] In some embodiments, the dynamic hypergraph diffusion branch includes a plurality of cascaded dynamic hypergraph diffusion unit groups; each dynamic hypergraph diffusion unit group includes two cascaded dynamic hypergraph diffusion units, the output features of the first dynamic hypergraph diffusion unit are input into the mutual attention operation, and the output features of the mutual attention operation are added element-wise to the output features of the second dynamic hypergraph diffusion unit.

[0025] The structure of each dynamic hypergraph diffusion unit is as Figure 2 shown. Among them, according to the input image or input features, the corresponding basic hypergraph adjacency matrix is obtained. Since there are a large number of mixed pixels in the low-resolution hyperspectral image, the components contained in the mixed pixels with similar spectra usually have a high degree of correlation. Therefore, the present invention constructs a basic hypergraph based on spectral clustering. Specifically, in some embodiments, the process of obtaining the basic hypergraph adjacency matrix in each dynamic hypergraph diffusion unit includes: aggregating the K nodes with the highest spectral similarity to each target node in the input image or input features to each hyperedge through the K-nearest neighbor algorithm; assigning weights to each node on the hyperedge through the negative exponential function to obtain the basic hypergraph adjacency matrix.

[0026] In the embodiments of the present invention, the process of assigning weights to each node on the hyperedge through the negative exponential function can be expressed by the following formula: ; where represents the Euclidean distance between the i-th node and the target node on the hyperedge, represents the average Euclidean distance between all nodes and the target node on the hyperedge; represents the weight value assigned to the i-th node on the hyperedge; The basic hypergraph adjacency matrix is obtained by the following formula: ; Wherein, represents the basic hypergraph adjacency matrix, represents the constructed basic hypergraph based on spectral clustering, where the basic hypergraph is composed of hyperedges, and the hyperedges are composed of the weights of each node, represents the degree of nodes in the hypergraph, represents the degree of edges in the hypergraph, and the degree of nodes and the degree of edges are calculated from the basic hypergraph . represents the matrix composed of the weight values , and it is initialized as the identity matrix.

[0027] Noise is added to the input image or input features to obtain the diffusion features at the initial time step. In the embodiments of the present invention, noise with a noise coefficient of is added to the input image or input features of the standard normal distribution noise, that is: ; Wherein, represents the standard normal distribution noise, that is , represents the noise coefficient, that is , represents the features after adding noise, and the features are the diffusion features at the initial time step, that is, the time step t = 0.

[0028] The corresponding time-varying feature vector is determined according to the number of channels of the input image or input features. In some embodiments, the process of determining the corresponding time-varying feature vector according to the number of channels of the input image or input features includes: The preliminary time-varying feature vector is generated by using sine time encoding through the following formula: ; ; Wherein, k represents the serial number of the frequency component, represents the fundamental frequency of the kth frequency component, d represents the number of channels of the input image or input features, and the number of channels of the input image is the number of spectral bands, represents the preliminary time-varying feature vector at the tth iteration. The preliminary time-varying feature vector is dimensionally compressed by multiple consecutive MLPs to obtain the time-varying feature vector.

[0029] In the embodiments of the present invention, for the preliminary time-varying feature vector Perform two consecutive MLP operations, and use the SiLU activation function to activate the features between the two MLP operations to obtain the time-varying feature vector at the t-th iteration. . Among them, the SiLU activation function can significantly improve the feature learning ability of the system through smooth nonlinearity and excellent gradient characteristics. The time-varying feature vector The acquisition process can be expressed as: ; Among them, Represents the weight when performing the first MLP operation When, Represents the weight when performing the second MLP operation When.

[0030] The time-varying feature vector is linearly mapped in combination with the diffusion feature of the current time step to obtain a linearly mapped feature. In the embodiment of the present invention, the time-varying feature vector Perform a broadcast operation to realize the automatic expansion of the dimension of the time-varying feature vector , and then broadcast the feature Is concatenated with the feature after adding noise To obtain the feature , that is: ; Among them, Represents the channel concatenation operation, Represents the diffusion feature of the current time step; The feature Is linearly mapped to a unified space and the feature channels are adjusted to obtain a linearly mapped feature , that is: ; Among them, Represents the weight of the linear mapping.

[0031] The embedding of the time vector is to regulate the inverse diffusion process. During the T-time step iteration, the embedding vector helps the model distinguish the noise feature distributions at different stages for the different noise intensities corresponding to each time step t. In the early stage, the noise variance is large, and the effective signal is concentrated in the low frequency. The time-varying feature vector Activates neurons responsible for low-frequency filtering; in the later stage, the high-frequency components are prominent, and the time-varying feature vector Activates neurons responsible for high-frequency filtering.

[0032] Perform a query-key attention operation on the linearly mapped feature to obtain a similarity matrix. In the embodiment of the present invention, the similarity matrix of the linearly mapped feature Of the current time step is observed through the query-key attention mechanism , that is: ; ; ; Among them, represents the query matrix obtained by the query-key attention operation, represents the query weight, represents the key matrix obtained by the query-key attention operation, represents the key weight, and C represents the number of channels of the linear mapping feature .

[0033] Then, calculate the Hadamard product of the similarity matrix and the corresponding basic hypergraph adjacency matrix to obtain the corresponding hypergraph adjacency matrix. In the embodiments of the present invention, the hypergraph adjacency matrix is obtained through the following formula: ; Among them, represents the hypergraph adjacency matrix, represents the Hadamard product operation.

[0034] Calculate the product of the corresponding elements of the hypergraph adjacency matrix, the linear mapping feature, and the learnable projection matrix to obtain the affine feature in the current dynamic hypergraph diffusion unit. In the embodiments of the present invention, the affine feature is obtained through the following formula: ; Among them, represents the affine feature, represents the learnable projection matrix, which is obtained by system initialization and belongs to the system learning parameters.

[0035] Denoise and extract features from the affine feature, and inject the processed features into the input image or input features. In some embodiments, the process of denoising and extracting features from the affine feature in each dynamic hypergraph diffusion unit includes: performing a convolution operation on the affine feature, performing a group normalization operation on the convolved feature, and performing a GELU activation operation on the feature after the group normalization operation to obtain the output feature. In the embodiments of the present invention, the process of denoising and extracting features from the affine feature to obtain the output feature can be represented by the following formula: ; Among them, Conv represents a two-dimensional convolution operation, GN represents a group normalization operation (GroupNorm), and group normalization can effectively eliminate the dependence on the batch size and enhance the task generalization.

[0036] In some embodiments, in each dynamic hypergraph diffusion unit, the process of injecting the processed features into the input image or input features to obtain the diffusion features at the next time step includes: ; ; ; wherein, represents the noise attenuation rate, represents the detail injection rate, and LinearSchedule represents a function for linearly adjusting hyperparameters, which is a tool for linearly adjusting hyperparameters in machine learning.

[0037] Repeating the processing of the diffusion features at the current time step for the diffusion features at the next time step to obtain the final diffusion features of the current dynamic hypergraph diffusion unit. In the embodiments of the present invention, for the diffusion features at the next time step repeating the above processing operation of the diffusion features at the current time step for T time steps of iteration to obtain the final diffusion features of the current dynamic hypergraph diffusion unit .

[0038] In some embodiments, in the differential-frequency collaborative attention branch, after performing convolutional feature extraction on the input image to obtain input features, the input features are processed by a plurality of cascaded differential-frequency collaborative attention unit groups; wherein, each differential-frequency collaborative attention unit group includes two cascaded differential-frequency collaborative attention units, the output features of the first differential-frequency collaborative attention unit are input into the mutual attention operation, and the output features of the mutual attention operation are added to the output features of the second differential-frequency collaborative attention unit element by element.

[0039] In the embodiments of the present invention, the input features are obtained by the following formula: ; wherein, represents the input features, I represents the input image, and ReLU represents the ReLU activation operation.

[0040] The differential-frequency collaborative attention branch includes two cascaded differential-frequency collaborative attention unit groups, each differential-frequency collaborative attention unit group includes two cascaded differential-frequency collaborative attention units, and the output features of the first differential-frequency collaborative attention unit are input into the mutual attention operation, and are combined with the output features of the first dynamic hypergraph diffusion unit and input into the mutual attention operation. It can be understood that the input of the mutual attention operation includes the collaborative features output by the differential-frequency collaborative attention branch and the diffusion features output by the dynamic hypergraph diffusion branch. In the embodiments of the present invention, the mutual attention operation is represented by the following formula: ; wherein, represents the output feature of the mutual attention operation, represents the query matrix of the diffusion feature, represents the key matrix of the collaborative feature, represents the value matrix of the collaborative feature, represents the mutual attention operation.

[0041] The structure of each differential-frequency collaborative attention unit is as Figure 3 shown, including a differential attention analysis sub-branch and a frequency analysis sub-branch. The input feature is processed by both the differential attention analysis sub-branch and the frequency analysis sub-branch to obtain the differential feature and the frequency feature respectively; the differential feature and the frequency feature are added element by element, and after the added feature is subjected to a convolution operation, it is added to the input feature element by element to obtain the collaborative feature.

[0042] In the embodiment of the present invention, residual processing is performed on the input feature. Specifically, the basic residual block in the classical ResNet network is used to perform residual processing on the input feature, and the feature after residual processing is added to the feature after addition and convolution element by element to obtain the collaborative feature.

[0043] The processing process of each differential-frequency collaborative attention unit can be represented by the following formula: ; wherein, represents the collaborative feature, represents the feature after residual processing of the input feature, represents the frequency feature, represents the differential feature.

[0044] In the differential attention analysis sub-branch, the input feature is convolved with multiple groups of spectral differential kernels of different scales to complete the differential calculation of the multi-scale spectral dimension sliding. The obtained multi-scale spectral features are processed by a multi-layer perceptron to generate multi-scale attention weights; the multi-scale attention weights are fused and then subjected to a convolution fine-tuning operation, and then multiplied by the input feature element by element to obtain the differential feature.

[0045] In the embodiment of the present invention, for the input feature one-dimensional convolution processing is performed using spectral differential kernels of [-1, 1, 0], [-1, 0, 1], and [-1, -1, 0, 1, 1] respectively, that is: ; ; ; wherein, , and respectively represent the spectral features after convolution with 3 groups of spectral difference kernels. represents the calculation of the difference by sliding along the spectral dimension; Perform channel average pooling on the spectral feature , spectral feature and spectral feature respectively to obtain the global statistics of the difference information, that is: ; where represents the j-th spectral feature after pooling, j = 1, 2, 3, and H and W respectively represent the height and width of the feature; Perform channel average pooling on the input feature as well to obtain the global statistics of the input feature , that is: ; where represents the spectral feature after pooling.

[0046] Generate multi-scale attention weights , , and corresponding to the multi-scale spectral features through the multi-layer perceptron MLP of the bottleneck structure, that is , , and , that is: ; After performing the fusion process of element-wise addition on the multi-scale attention weights, then perform the fine-tuning operation of 1×1 convolution, that is ; where represents the fused weight, represents the Sigmoid activation operation, represents the 2D convolution operation of 1×1; Multiply the weight element-wise with the input feature to obtain the difference feature , that is: .

[0047] In the frequency analysis sub-branch, the Laplacian operator and the Gaussian kernel are used to extract high-frequency features and low-frequency features from the input features respectively; after channel concatenation of the high-frequency features and the low-frequency features, a convolution fusion operation is performed on the concatenated features, and then the fused features are multiplied by the corresponding elements of the input features to obtain the frequency features.

[0048] In the embodiment of the present invention, the Laplacian operator and the Gaussian kernel are used to extract from the input features high-frequency features and low-frequency features respectively, that is: ; wherein, Laplacian represents the Laplacian operator, Gaussian represents the Gaussian kernel, represents a 3×3 two-dimensional convolution operation; After channel concatenation of the high-frequency features and the low-frequency features , a convolution fusion operation is performed on the concatenated features, and then the fused features are multiplied by the corresponding elements of the input features to obtain the frequency features , that is: .

[0049] In some embodiments, as Figure 1 shown, in the super-resolution reconstruction sub-branch, the diffusion features and the collaborative features are added corresponding elements, and then the added features are continuously subjected to convolution operations and transposed convolution operations, and then added to the upsampled image corresponding elements; after performing a convolution operation on the added features, a super-resolution image is obtained.

[0050] In the super-resolution reconstruction sub-branch of the embodiment of the present invention, the diffusion features and the collaborative features are added corresponding elements, and then the added features are continuously subjected to convolution operations and transposed convolution operations, that is: ; wherein, represents a transposed convolution operation, represents the currently obtained features; The features are added to the bicubic upsampled image of the input image corresponding elements, and after performing a convolution operation on the added features, a super-resolution image is obtained, that is: ; wherein, Represents the bicubic upsampled image of the input image.

[0051] The present invention also provides a super-resolution method for remotely sensed hyperspectral images with chain diffusion, as Figure 4 shown. The method includes: S1: Obtain a remotely sensed image dataset, and preprocess the remotely sensed image dataset to obtain a training set.

[0052] In the embodiments of the present invention, the process of preprocessing the remotely sensed image dataset includes: segmenting the remotely sensed hyperspectral image hyperspectral dataset, selecting 70% of the image area as the training part in the training set, and 30% of the area as the validation part in the training set. Randomly select patches from each area of the remotely sensed image, set different patch sizes according to the magnification factor, where the 2x, 3x, and 4x magnifications correspond to patch sizes of 64×64, 96×96, and 128×128 respectively; perform random horizontal flipping, rotation at different angles, and scaling at different ratios on each patch. Finally, downsample these patches to 32×32 low-resolution hyperspectral images according to different scale factors through bicubic downsampling to obtain the training set.

[0053] S2: Input the training set obtained in step S1 into the super-resolution system for remotely sensed hyperspectral images with chain diffusion provided by the present invention, and train the super-resolution system in combination with the semantic constraint loss function to obtain a super-resolution model.

[0054] During the training process, the training part in the training set completes the training of the super-resolution system, and the validation part in the training set completes the calculation and verification of the semantic constraint loss function for the currently trained model, and feeds back to the super-resolution model based on the semantic constraint loss function to adjust the weight parameters in the super-resolution model.

[0055] In some embodiments, the semantic constraint loss function is obtained by the following formula: ; where represents the semantic constraint loss function, represents the L1 loss function, represents the SAM loss function, represents the superpixel nesting loss function, , and represent loss weights, and the weights are adaptively adjusted according to the actual training situation. In the embodiments of the present invention, the weights , and are 1, 0.2, and 0.2 respectively The superpixel nesting loss function Obtained from the following formula: ; wherein, represents the spectral loss, represents the texture loss, represents the superpixel nesting loss weight, which is adaptively adjusted according to the actual training situation; Spectral loss Obtained from the following formula: ; wherein, represents the number of superpixels obtained after superpixel segmentation of the super-resolution image output by the super-resolution model using the SLIC algorithm, represents the spectral mean of all pixels within the i-th superpixel in the corresponding true high-resolution remote sensing image of the super-resolution image, represents the spectral mean of all pixels within the i-th superpixel in the super-resolution image, and the spectral mean is obtained by the following formula: ; wherein, represents the superpixel, represents the number of pixels included in the superpixel, represents the pixel value at the (x, y) position in image I; Texture loss Obtained from the following formula: ; wherein, represents the texture entropy of all pixels within the i-th superpixel in the true high-resolution remote sensing image, represents the texture entropy of all pixels within the i-th superpixel in the super-resolution image, and the texture entropy is obtained by the following formula: ; wherein, represents the probability of the b-th bin in the histogram distribution of LBP values within the superpixel, and B represents the number of intervals of the histogram.

[0056] S3: According to the super-resolution model obtained in step S2, adjust the hyperparameters of the training super-resolution system, and repeat step S2 until the optimal super-resolution model is obtained.

[0057] In the embodiment of the present invention, the optimizer is selected as Adam, where the first-moment exponential decay factor and the second-moment exponential decay factor , the batch size is 4 and the initial learning rate is 0.0001. The number of training epochs is set to 40. The model parameters are updated according to the semantic constraint loss function. As the semantic constraint loss function gradually decreases, the super-resolution reconstruction effect of the model on low-resolution images becomes better and better. After several rounds of updating and iterating the network parameters, the performance of the model is tested using the validation part of the training set. When the results of the objective image quality evaluation metrics of the reconstructed images in the validation part no longer increase or reach the preset upper limit of the number of training epochs, the training ends.

[0058] S4: Input the low-resolution remote sensing image into the optimal super-resolution model obtained in step S3 for super-resolution reconstruction to obtain the corresponding high-resolution remote sensing image.

[0059] To clearly demonstrate the excellent super-resolution reconstruction effect of the chain diffusion-based remote sensing hyperspectral image super-resolution system and method provided by the present invention, the embodiments of the present invention use peak signal-to-noise ratio (PSNR), structural similarity (SSIM), spectral angle similarity (SAM), and relative global error separation metric (ERGAS) to objectively evaluate the quality of the reconstructed remote sensing hyperspectral image, and use bicubic interpolation (Bicubic), 3D-FCNN (published in the paper "Hyperspectral Image Spatial Super-Resolution via 3D Full Convolutional Neural Network" in the journal "Remote Sensing"), GDRRN (published in the paper "Single Hyperspectral Image Super-resolution with Grouped Deep Recursive Residual Network" in the journal "2018 IEEE Fourth International Conference on Multimedia Big Data (BigMM)"), SSPSR (published in the paper "Learning Spatial-Spectral Prior for Super-Resolution of Hyperspectral Imagery" in the journal "IEEE TRANSACTIONS ON COMPUTATIONAL IMAGING"), MCNet (published in the paper "Mixed 2D / 3D Convolutional Network for Hyperspectral Image Super-Resolution" in the journal "Remote Sensing"), ERCSR (published in the paper "Exploring the Relationship Between 2D / 3D Convolution for Hyperspectral Image Super-Resolution" in the journal "IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING"), GELIN (published in the paper "A Group-Based Embedding Learning and Integration Network for Hyperspectral ImageSuper-Resolution》), SRDNet (the paper "Hyperspectral Image Super-Resolution via Dual-Domain Network Based on Hybrid Convolution" published in the journal IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING), and MSSR (the paper "Remote Sensing Hyperspectral Image Super-Resolution via Multidomain Spatial Information and Multiscale Spectral Information Fusion" published in the journal IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING), the CST method (the paper "Cross-Scope Spatial-Spectral Information Aggregation for Hyperspectral Image Super-Resolution" published in the journal IEEE TRANSACTIONS IMAGE PROCESS) and the method provided by the present invention (CDiff-HG) are respectively in the MDAS dataset, PaviaThe Centre dataset and the Houston dataset are compared. The objective comparison results of the method provided by the present invention and the comparative method on the MDAS dataset are shown in Table 1. The present invention has achieved the optimal comprehensive performance under all scaling factors, verifying its effectiveness in the joint modeling of spatial-spectral features. From the perspective of objective indicators, the present invention performs outstandingly in two key indicators of PSNR and SSIM. Taking the ×2 super-resolution reconstruction as an example, the PSNR of the present invention reaches 34.487 dB, which is 0.165 dB higher than that of the sub-optimal model MSSR (34.322 dB); the SSIM is 0.9244, which is better than 0.9215 of MSSR. This advantage is still significant at higher scaling factors (×3, ×4), indicating that the model has a strong ability to recover high-frequency details. In addition, in terms of the spectral fidelity index SAM and the error sensitivity index ERGAS, the present invention also performs excellently, verifying its ability to suppress spectral distortion. Compared with the existing methods, the PSNR of the present invention reaches 31.576 dB in the ×3 super-resolution task, which is 0.146 dB and 0.125 dB higher than that of SRDNet (31.430 dB) and MSSR (31.451 dB) respectively, and the SSIM (0.8524) is significantly better than other models. It should be noted that in the ×4 super-resolution, although the increase in PSNR is small, SAM and ERGAS still remain at the highest level. This indicates that while maintaining spectral continuity, the model has more precise control over the reconstruction error. This characteristic is particularly important for subsequent interpretation tasks of hyperspectral images. Although the existing methods have improved performance through 3D convolution or Transformer structures, their modeling of the global trend and local morphology in the spectral dimension still has limitations. The present invention explicitly models the local correlation and global distribution characteristics of spatial-spectral features in hyperspectral images by introducing dynamic hypergraph learning and diffusion models, thus showing stronger robustness in complex degradation scenarios. Figure 5 The visualization results in the ×4 super-resolution task are shown. Generally speaking, the images reconstructed by the present invention have clearer contours and show more detailed information expression in areas with complex textures. Taking the annotated areas of the visualization images as an example, the houses and roads restored by the present invention have more realistic structures and less blur.

[0060] Table 1:

[0061] The objective comparison results of the method provided by the present invention and the comparative method on the Pavia Centre dataset are shown in Table 2. The present invention shows significant advantages under all scaling factors, especially in the balance between spatial detail restoration and spectral fidelity. In the ×2 super-resolution task, the PSNR (36.376 dB) and SSIM (0.9565) of the present invention both reach the optimum, improving by 0.140 dB and 0.0013 respectively compared with MSSR (36.236 dB, 0.9552). At the same time, its SAM and ERGAS are also the lowest values, indicating that the model effectively suppresses spectral distortion while improving the spatial resolution. It is worth noting that when the scaling factor increases to ×3, the PSNR (31.807 dB) and SSIM (0.8896) of the present invention are more significantly improved compared with MSSR (31.547 dB, 0.8836), reaching 0.260 dB and 0.0060 respectively, further verifying the robustness of the method. In the challenging ×4 super-resolution task, the PSNR and SSIM of the present invention both exceed all comparative models. Among them, the PSNR is improved by 0.117 dB compared with the sub-optimal model SSPSR, and the SSIM is improved by 0.0041. The present invention constructs a spatial-spectral joint representation through a hypergraph diffusion unit and a differential-frequency collaborative attention mechanism, combined with a superpixel nested loss function, which can adaptively capture local structural features and global distribution laws at different scales. The performance advantages of the proposed method on the Pavia Centre dataset further verify its generalization ability across datasets. Figure 6 Figure 6 shows the visualization results in the ×4 super-resolution task. Compared with earlier models, the images reconstructed by MSSR, CST, and the present invention have richer textures. In the annotated area, only the present invention can restore the edge shape of the original building to the greatest extent.

[0062] Table 2:

[0063] The objective comparison results of the method provided by the present invention and the comparative method on the Houston dataset are shown in Table 3. The present invention has achieved the best results in the ×2, ×3, and ×4 super-resolution tasks, further verifying its generalization ability and stability in different scenarios. In the ×2 super-resolution task, the PSNR (38.226 dB) and SSIM (0.9506) of the present invention are significantly better than those of other models, increasing by 0.139 dB and 0.0016 respectively compared to MSSR (38.087 dB, 0.9490). At the same time, its SAM (2.520) and ERGAS (5.955) reach the lowest values, indicating that the model can effectively enhance spatial details and maintain spectral fidelity. Compared with the Transformer-based CST model, the PSNR of the present invention increases by 0.323 dB and the SAM decreases by 4.26%, verifying the advantages of dynamic hypergraph learning in spectral correlation modeling. Similarly, for the ×3 and ×4 super-resolution tasks, the present invention still has a stable leading advantage. This shows that its progressive optimization strategy can more effectively suppress the accumulation of reconstruction errors, highlighting the strong robustness of the model in spectral-spatial reconstruction at different magnification scales. Figure 7 The visualization results in the ×4 super-resolution task are shown. Generally speaking, CST and the present invention seem to have the best performance, especially in effectively alleviating the blurring effect. In the annotated area of the visualization results, only the present invention restores coherent and accurately angled edges compared to other methods.

[0064] Table 3:

[0065] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the disclosure of the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in the present invention can be achieved. No limitations are imposed herein.

[0066] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A remote sensing hyperspectral image super-resolution system with chain diffusion, characterized in that, Comprising: A dynamic hypergraph diffusion branch, which embeds the time condition of the input image into the input image, updates and learns the dynamic adjacency matrix for the obtained features, and obtains affine features based on the dynamic adjacency matrix; denoises and extracts details from the affine features to obtain diffusion features; A differential-frequency collaborative attention branch, which performs a two-branch spatial attention operation on the input image, completes differential and frequency analysis, and respectively obtains differential features and frequency features; Combining the differential features and the frequency features to obtain collaborative features; A super-resolution reconstruction branch, which combines the diffusion features, the collaborative features, and the upsampled image of the input image to obtain a super-resolution image; Feature fusion is performed through mutual attention operation between the dynamic hypergraph diffusion branch and the differential-frequency collaborative attention branch, and the fused features are used to respectively guide the dynamic hypergraph diffusion branch and the differential-frequency collaborative attention branch for feature processing.

2. The chain diffusion-based remote sensing hyperspectral image super-resolution system according to claim 1, wherein The dynamic hypergraph diffusion branch includes multiple cascaded dynamic hypergraph diffusion unit groups; each dynamic hypergraph diffusion unit group includes two cascaded dynamic hypergraph diffusion units, the output features of the first dynamic hypergraph diffusion unit are input into the mutual attention operation, and the output features of the mutual attention operation are added element-wise to the output features of the second dynamic hypergraph diffusion unit; In each dynamic hypergraph diffusion unit: According to the input image or input features, obtain the corresponding basic hypergraph adjacency matrix; Add noise to the input image or input features to obtain the diffusion features at the initial time step, and at the same time determine the corresponding time-varying feature vector according to the number of channels of the input image or input features; linearly map the time-varying feature vector combined with the diffusion features at the current time step to obtain linearly mapped features; Perform a query-key attention operation on the linearly mapped features to obtain a similarity matrix; Then calculate the Hadamard product of the similarity matrix and the corresponding basic hypergraph adjacency matrix to obtain the corresponding hypergraph adjacency matrix; Calculate the element-wise product of the hypergraph adjacency matrix, the linearly mapped features, and the learnable projection matrix to obtain the affine features in the current dynamic hypergraph diffusion unit; Denoise and extract features from the affine features, and inject the processed features into the input image or input features to obtain the diffusion features at the next time step; Repeat the processing of the diffusion features at the current time step for the diffusion features at the next time step, and perform multiple time step iterations to obtain the final diffusion features of the current dynamic hypergraph diffusion unit.

3. The chain diffusion-based remote sensing hyperspectral image super-resolution system according to claim 2, characterized in that, The process of obtaining the basic hypergraph adjacency matrix in each dynamic hypergraph diffusion unit includes: Through the K-nearest neighbor algorithm, aggregate the K nodes with the highest spectral similarity to each target node in the input image or input features to each hyperedge; Assign weights to each node on the hyperedge through the negative exponential function to obtain the basic hypergraph adjacency matrix.

4. The chain diffusion-based remote sensing hyperspectral image super-resolution system according to claim 2, wherein The process of determining the corresponding time-varying feature vector according to the number of channels of the input image or input features includes: Use the following formula to generate a preliminary time-varying feature vector through sinusoidal time encoding: ; ; where k represents the serial number of the frequency component, denotes the fundamental frequency of the k-th frequency component, d represents the number of channels of the input image or input feature, represents the preliminary time-varying feature vector at the t-th iteration; Perform dimensional compression on the preliminary time-varying feature vector through multiple consecutive MLPs to obtain the time-varying feature vector.

5. The chain diffusion-based remote sensing hyperspectral image super-resolution system according to claim 2, wherein The process of denoising and feature extraction for the affine features in each dynamic hypergraph diffusion unit includes: performing a convolution operation on the affine features, performing a group normalization operation on the convolved features, and performing a GELU activation operation on the features after the group normalization operation to obtain output features.

6. The chain diffusion-based remote sensing hyperspectral image super-resolution system according to claim 2, characterized in that, The process of injecting the processed features into the input image or input features in each dynamic hypergraph diffusion unit includes: ; ; ; Among them, represents the noise attenuation rate, represents the detail injection rate, represents the input image or the current input feature, represents the feature after denoising and feature extraction processing of the affine feature, and LinearSchedule represents a function for linearly adjusting hyperparameters.

7. The chain diffusion-based remote sensing hyperspectral image super-resolution system according to claim 1, characterized in that, In the differential-frequency collaborative attention branch, after convolutional feature extraction is performed on the input image to obtain input features, the input features are processed by multiple cascaded groups of differential-frequency collaborative attention units; where each group of differential-frequency collaborative attention units includes two cascaded differential-frequency collaborative attention units, the output features of the first differential-frequency collaborative attention unit are input into the mutual attention operation, and the output features of the mutual attention operation are added element-wise to the output features of the second differential-frequency collaborative attention unit; each differential-frequency collaborative attention unit includes a differential attention analysis sub-branch and a frequency analysis sub-branch, and the input features are processed by the differential attention analysis sub-branch and the frequency analysis sub-branch respectively to obtain the differential features and the frequency features correspondingly; the differential features and the frequency features are added element-wise, and after the added features are subjected to a convolution operation, they are added element-wise to the input features to obtain the collaborative features; In the differential attention analysis sub-branch, the input features are convolved with multiple groups of spectral difference kernels of different scales to complete multi-scale spectral dimension sliding calculation of differences; the obtained multi-scale spectral features are processed by a multi-layer perceptron to generate multi-scale attention weights; after the multi-scale attention weights are fused, a convolution fine-tuning operation is performed, and then they are multiplied element-wise with the input features to obtain the differential features; In the frequency analysis sub-branch, a Laplace operator and a Gaussian kernel are used to extract high-frequency features and low-frequency features from the input features respectively; after the high-frequency features and the low-frequency features are concatenated in channels, a convolution fusion operation is performed on the concatenated features, and then the fused features are multiplied element-wise with the input features to obtain the frequency features.

8. The remote sensing hyperspectral image super-resolution system with chain diffusion according to claim 1, characterized in that, In the super-resolution reconstruction branch, the diffusion features and the collaborative features are added element-wise, and then after the added features are continuously subjected to a convolution operation and a transposed convolution operation, they are added element-wise to the upsampled image; after the added features are subjected to a convolution operation, the super-resolution image is obtained.

9. A super-resolution method for remote sensing hyperspectral images with chain diffusion, characterized in that Including: S1: Obtain a remote sensing image dataset, and preprocess the remote sensing image dataset to obtain a training set; S2: Input the training set obtained in step S1 into the chain diffusion-based remote sensing hyperspectral image super-resolution system according to any one of claims 1 to 8, and train the super-resolution system in combination with a semantic constraint loss function to obtain a super-resolution model; S3: According to the super-resolution model obtained in step S2, adjust and train the hyperparameters of the super-resolution system, and repeat step S2 until an optimal super-resolution model is obtained; S4: Input the low-resolution remote sensing image into the optimal super-resolution model obtained in step S3 for super-resolution reconstruction to obtain the corresponding high-resolution remote sensing image.

10. The chain diffusion-based remote sensing hyperspectral image super-resolution method according to claim 9, characterized in that The semantic constraint loss function in step S2 is obtained by the following formula: ; Among them, represents the semantic constraint loss function, represents the L1 loss function, represents the SAM loss function, represents the superpixel nesting loss function, , and represent the loss weights; Superpixel nested loss function Obtained by the following formula: ; Among them, represents the spectral loss, represents the texture loss, represents the superpixel nesting loss weight; Spectral loss Obtained from the following formula: ; Among them, represents the number of superpixels obtained after superpixel segmentation of the super-resolution image output by the super-resolution model using the SLIC algorithm, represents the spectral mean of all pixels within the i-th superpixel in the corresponding true high-resolution remote sensing image of the super-resolution image, represents the spectral mean of all pixels within the i-th superpixel in the super-resolution image. The spectral mean is obtained by the following formula: ; Among them, represents a superpixel, represents the number of pixels included in the superpixel, represents the pixel value at the position (x, y) in the image I; Texture loss Obtained from the following formula: ; Among them, represents the texture entropy of all pixels within the i-th superpixel in the true high-resolution remote sensing image, represents the texture entropy of all pixels within the i-th superpixel in the super-resolution image, and the texture entropy is obtained by the following formula: ; wherein, represents the probability of the b-th bin in the histogram distribution of LBP values within a superpixel, and B represents the number of intervals of the histogram.

Citation Information

Patent Citations

  • Location recommendation method based on hypergraph neural network and diffusion model

    CN119441636A

  • Transform-based remote sensing image super-resolution reconstruction method

    CN119887525A

  • Remote sensing hyperspectral image super-resolution reconstruction method based on hybrid neural network

    CN119904359A

  • Image processing device, image processing method and medium

    US20150317767A1

  • Method and apparatus with frame image reconstruction

    US20240193727A1

Cited By

  • SMT welding spot detection method based on hypergraph contrast learning

    CN120931575A

  • A hypergraph contrastive learning based SMT solder joint detection method

    CN120931575B

  • Power inspection video diffusion super-division cooperative enhancement method for equipment defect identification

    CN121860874A

  • Lightweight image super-resolution method based on super-pixel guidance

    CN121903843A

  • A lightweight image super-resolution method based on superpixel guidance

    CN121903843B